Should a small language model stay on the device for a specific job, or should the app call a server? Consider a phone app that turns a meeting note into three action items: owner, action and due date. It must refuse to invent a date. This is narrower than an assistant that researches the meeting topic. Local inference could avoid transmitting the note, but the answer depends on the actual app, supported devices and network policy—not the size printed on a model card.
Decide the job before choosing the model
Set an output schema and a stop rule: if an owner or deadline is absent, return “not specified”; never send an invite without confirmation. Compare a deterministic extractor, a compact local model, a server model and a hybrid router on the same held-out notes. A rules-based path may win on this constrained task. Meta's Llama 3.2 announcement describes 1B and 3B text models for edge tasks such as summarization and rewriting; that establishes availability, not fitness for this workflow. Platform availability is another constraint: Google's LLM Inference guide marks its Android and iOS implementations deprecated (not the web implementation) and directs mobile projects toward LiteRT-LM. Pin a runtime and model version rather than designing around a marketing category.
Use local only if it passes every hard gate, including the offline path. Use server if permitted egress and the measured end-to-end benefit justify it. Hybrid is not a free compromise: its router, user consent and offline failure behavior are part of the product. If the workflow processes regulated or particularly sensitive text, obtain a legal/security review before choosing any path; this article is not a compliance determination.
A protocol another team can repeat
- Freeze the test. Collect consented, de-identified notes from the intended workflow, with ordinary, long, multilingual, OCR-corrupted, ambiguous, contradictory and adversarial examples. Separate a held-out set; version it and predeclare acceptable quality by stratum, including the correct refusal behavior. Have two reviewers adjudicate disagreements against the original note, not a model's fluent phrasing.
- Hold the comparison constant. Record app build, OS, device class (including the least capable supported one), runtime/backend, model and quantization hashes, prompt/schema, tokenizer, input/output lengths, server region, network type and concurrency. Repeat across at least two charge and thermal states. Randomize candidate order and report distributions and confidence intervals rather than one fastest run. An on-device inference study evaluates multiple models, quantizations and device configurations: its results are evidence that configurations matter, not transferable phone benchmarks.
- Measure the experience. For a fresh process and a warmed session separately, log start-to-first-useful-output and completion latency (p50/p95), quality by stratum, refusal/error rate, peak app working set alongside other apps, storage/download size and incremental battery energy per completed task. Include preprocessing, model load, queueing, upload, routing and postprocessing; compare against the existing workflow. Apple's battery analysis guidance recommends Xcode, MetricKit and Instruments and notes temperature-related limits; use platform instruments and a controlled baseline, not token/s as an energy proxy.
- Break the route. Test airplane mode, captive portal, slow and intermittent network, server timeout, unavailable model download, insufficient storage, memory pressure, revoked consent, old OS, and a model update. Record whether output is correct, delayed, refused or silently rerouted. Inspect network telemetry and crash payloads to verify which raw fields actually leave the device.
For each candidate, preserve per-item outcomes and failure categories. If no route meets the gates, do not select a winner by averaging: narrow the task, add human review or keep the current process. Rerun the same suite after quantization, prompt changes and runtime updates. A cold-start penalty can reverse a warm-demo ranking; an offline requirement can disqualify an otherwise superior remote path.
A deliberately synthetic decision gate
The following Python 3 standard-library program runs without a model or network. Every number in its rows and thresholds is illustrative, invented for policy testing—not a device benchmark or recommendation. Save the code as decision_harness.py and run python decision_harness.py, then python decision_harness.py --allow-egress --no-offline. Replace the rows with measurements from your own application; keep the policy thresholds agreed before looking at results. “Monthly” is a fabricated comparable operating-cost estimate, not a complete TCO calculation. Device energy excludes server power.
#!/usr/bin/env python
"""Illustrative decision gate, NOT a model or device benchmark. Python 3 stdlib.
Usage: python decision_harness.py [--allow-egress] [--no-offline]
Replace ROWS with measured values from your own application and devices.
"""
import argparse
# Each row represents one workload stratum, with cold/warm application completion
# latency (seconds), measured peak working set (MiB), incremental device energy
# (joules/request), independently adjudicated task success fraction, and flags.
# All numbers below are fabricated to exercise the policy; do not quote them as
# hardware performance or infer server-side energy from device energy.
ROWS = [
dict(route="local", stratum="routine", quality=.97, cold=2.8, warm=.8,
memory=740, energy=3.2, raw_exits=False, offline=True, monthly=95),
dict(route="local", stratum="hard", quality=.77, cold=4.6, warm=1.8,
memory=790, energy=5.1, raw_exits=False, offline=True, monthly=95),
dict(route="hybrid", stratum="routine", quality=.97, cold=3.0, warm=.9,
memory=760, energy=3.4, raw_exits=False, offline=True, monthly=110),
dict(route="hybrid", stratum="hard", quality=.94, cold=5.1, warm=2.4,
memory=760, energy=4.6, raw_exits=True, offline=False, monthly=110),
dict(route="server", stratum="routine", quality=.98, cold=2.1, warm=1.2,
memory=130, energy=1.0, raw_exits=True, offline=False, monthly=140),
dict(route="server", stratum="hard", quality=.96, cold=3.8, warm=2.7,
memory=130, energy=1.5, raw_exits=True, offline=False, monthly=140),
]
# Example policy only: substitute thresholds, strata and costs before a real decision.
QUALITY_FLOOR = {"routine": .90, "hard": .90}
MAX_COLD_S, MAX_WARM_S = 6.0, 3.0
MAX_MEMORY_MIB, MAX_DEVICE_ENERGY_J = 900, 6.0
def assess(rows, allow_egress=False, require_offline=True):
by_route = {}
for row in rows:
by_route.setdefault(row["route"], []).append(row)
results = {}
for route, entries in sorted(by_route.items()):
issues = []
strata = [r["stratum"] for r in entries]
if sorted(strata) != sorted(QUALITY_FLOOR):
issues.append("missing/duplicate strata")
if len({r["monthly"] for r in entries}) != 1:
issues.append("inconsistent monthly cost")
for r in entries:
label = r["stratum"]
if label not in QUALITY_FLOOR or r["quality"] < QUALITY_FLOOR.get(label, 1):
issues.append(label + ": quality")
if r["cold"] > MAX_COLD_S or r["warm"] > MAX_WARM_S:
issues.append(label + ": latency")
if r["memory"] > MAX_MEMORY_MIB:
issues.append(label + ": memory")
if r["energy"] > MAX_DEVICE_ENERGY_J:
issues.append(label + ": device energy")
if r["raw_exits"] and not allow_egress:
issues.append(label + ": raw data egress")
if not r["offline"] and require_offline:
issues.append(label + ": offline failure")
results[route] = (entries[0]["monthly"], issues)
eligible = [(cost, route) for route, (cost, issues) in results.items() if not issues]
return results, min(eligible)[1] if eligible else None
def main():
p = argparse.ArgumentParser(description=__doc__)
p.add_argument("--allow-egress", action="store_true")
p.add_argument("--no-offline", action="store_true")
args = p.parse_args()
results, winner = assess(ROWS, args.allow_egress, not args.no_offline)
for route, (cost, issues) in sorted(results.items()):
print(f"{route}: ${cost}/month; " + ("PASS" if not issues else "FAIL: " + ", ".join(issues)))
print("Decision:", winner or "no eligible route; redesign, relax policy explicitly, or stop")
if __name__ == "__main__":
main()
In the default example, local fails hard-case quality, hybrid fails offline and raw-egress gates, and server fails those two gates: the decision is no eligible route. With egress permitted and offline not required, hybrid passes and has lower illustrative monthly cost than server. That result is conditional, not a reason to turn off either requirement. The simple gate uses one aggregate quality fraction per stratum and one observation per thermal state; real data need uncertainty intervals, p95 tails, privacy audit evidence, an explicit sampling mix and a sensitivity analysis of costs and thresholds.
Account for what the handset and organization pay
Weight file size is not peak app memory: tokenizer, KV cache, runtime buffers, input length and competing apps all occupy resources. Budget download/storage and first-run loading separately from steady inference. Quantization may change both size and answer quality. If a server is used, price per request includes input/output volume, retries, region, retention and support; a hybrid adds router development, validation and monitoring, while local adds cross-device QA, distribution, updates and user battery cost. Compare cost at the expected request mix and at a high-usage case; don't treat a handset's energy draw as the entire system's footprint.
A local route can fail by returning a plausible invented deadline; a remote route can fail by transmitting a private note; a hybrid can fail both ways if its confidence router misclassifies a hard note or quietly falls back during a timeout. Define route-specific stop states. Keep a visible “on device” / “send for enhanced processing” indicator and require affirmative permission for the latter if the task promises local handling. Offline should mean no remote attempt, not an indefinitely spinning fallback. The related Meydo Journal discussion of phone hardware addresses the broader device constraints; this test is about a bounded workflow.
The security boundary travels with the data
Local inference reduces one transfer path but does not make a shared, lost or compromised device safe. Inventory local note caches, model packages, logs, telemetry, backup and crash reporting; minimize sensitive data and set retention/deletion behavior. A hybrid's router must decide before uploading text, and logs must show which route was used without retaining the note unnecessarily. Review the platform-specific privacy assurances rather than borrowing them wholesale: Google's AICore documentation describes request isolation and no retained input/output for its own processing, not a guarantee about an unrelated app's analytics or server fallback.
Pin and integrity-check model and runtime artifacts; distribute through a trusted update channel, stage rollout, retain a tested rollback, and rerun quality and privacy checks after each change. A compromised package or prompt-injected note can cross boundaries even when no cloud call occurs. Treat extracted action items as untrusted suggestions and require human approval before a calendar invitation or other external action. The release decision is therefore operational: ship local only if the measured least-capable-device experience and offline refusal behavior clear the agreed gates; otherwise choose an explicitly consented hybrid/server route if allowed—or decline automation for this task.
