Why Your On-Device Foundation Models App Feels Dumber Than the Demo

NativeFirst Team 7 min read
Macro photo of a computer chip circuit board

I was demoing ThinkBud on my own iPhone 17 Pro last week, feeling pretty smug about it, when a friend tried the exact same prompt on his iPhone 15. Different phrasing, weaker answer, same model, same app, same on-device Foundation Models framework doing the same job it always does.

Nothing crashed. Nothing errored. It just felt a little dumber on his phone than mine.

I chalked it up to prompt luck for about a day. Then I read a forum post that had nothing to do with Apple, iOS, or Swift, and it explained exactly what I’d seen.


The post that ruined my Saturday morning

It’s from Level1Techs, and it’s about running local LLMs on GPU clusters with vLLM — about as far from an iPhone as you can get. The author, going by thr3e, set out to answer a question every local-model tinkerer has asked: why does the model that benchmarks so well online feel worse when you run it yourself?

The answer isn’t “your GPU is bad” or “you picked the wrong quant.” It’s stranger than that. The same weights, run through different attention backends on the same GPU, produce measurably different outputs. Not different vibes — different tokens, provably, at the logit level.

The setup: a dense 27B model, full BF16 precision, no quantization anywhere, one variable changed — which of three available attention kernels (FlashAttention 2, Flash Infer, Triton) handled the math during prompt processing. Everything else pinned identical.

For the first few thousand tokens of a real ~100k-token workload, all three backends agreed on every single next-token prediction. Then, in later stretches of the prompt, they started disagreeing — sometimes picking a genuinely different greedy token, not just a slightly different probability.

Same model. Same GPU. Same weights. Different answer, depending on which kernel did the matrix multiplication.

If that doesn’t unsettle you a little, reread it. It unsettled me.


”Your implementation sucks. So does everyone else’s.”

That’s a direct quote, and it’s the actual thesis. Every LLM you run — cloud or local — sits on top of a stack nobody fully audits: attention kernels, quantization schemes, sampler settings, KV-cache precision, hardware instruction sets. Each layer can introduce tiny numerical divergence from whatever “reference” answer the model card benchmarked against.

The reference implementation — the one the lab used to publish its benchmark numbers — ran on hardware and software you will never touch. Yours runs on different hardware, a different framework, maybe a different quant. The gap between “benchmarked” and “what you experience” isn’t a bug in your setup. It’s the entire stack, compounding small differences until they occasionally flip which word comes next.

Now swap “vLLM on an RTX Pro 6000” for “Foundation Models framework on an A17 Pro Neural Engine,” and tell me that doesn’t describe your app too.


Apple hides the stack. It doesn’t remove it.

Here’s the thing that makes Apple’s on-device story feel safer than it is: you never see the kernel selection, the quantization scheme, or the sampler settings. SystemLanguageModel is one call. No backend flags, no precision knobs, no visibility into what actually ran.

That’s a genuine strength for shipping fast. It’s also exactly the kind of black box the Level1Techs post is warning you about — you’ve just outsourced the auditing to Apple instead of doing it yourself.

And the hardware really does differ under the hood. The Neural Engine has changed generation to generation — different core counts, different supported precisions, different scheduling behavior between an A17 Pro, an A18, and an M-series chip. Apple doesn’t publish which on-device configuration runs on which device for Foundation Models the way a GPU vendor publishes SM compute capability, but there’s no reason to assume the inference path is bit-identical across three chip generations spanning three years. It almost certainly isn’t.

I tested this exact framework back in July, wiring mi12labs/SwiftAI into a scratch project to try Apple’s on-device model with structured output. I asked it a simple factual question and got back a well-typed, perfectly structured answer that was also wrong:

The population, though: Ljubljana is a bit under 300,000 people, not 201,600. Apple’s on-device model got the shape of the answer right and the substance of the answer wrong by about a third.

At the time I filed that under “small models don’t know facts, that’s expected.” Reading the Level1Techs piece, I’d add a second, less comfortable line: I have no idea whether that specific wrong answer was a property of the model’s weights, or a property of my specific device’s inference path that day. Neither do you, for your app. Neither does Apple say.


The part that actually matters for shipping

You can’t get vLLM-style backend flags out of Foundation Models, and you shouldn’t try — that’s not the layer you control. But three things from that post translate directly:

Test on real hardware you don’t own. Simulator behavior and your personal iPhone are one data point each. If your app leans on Foundation Models for anything beyond decoration, borrow an older device — a genuinely different Neural Engine generation, not just a smaller screen — before you ship. “Works on my phone” means less here than it does for UI bugs.

@Generable fixes the shape, not the substance. I said this in July and the Level1Techs post is the mathematical version of the same point: structured output constrains what kind of answer you get, not whether the underlying inference converged on the right one. Don’t let a clean Codable struct talk you into trusting the values inside it.

Don’t zero-shot your own confidence. The post calls out testing a model with three prompts at temperature zero and calling it good or bad — “not a good analog of most agentic tasks.” Same trap in an app: if you tried a feature five times during development and it worked, that’s not evidence it works. Test the boring, repetitive, real-workload version, the same way you’d load-test an API instead of curling it once.


The honest takeaway

Nobody’s local LLM setup — or on-device iOS deployment — is broken. That’s the actual, slightly grim reassurance in the original post: everyone’s stack diverges a little from the reference. The mistake is assuming your stack is the exception.

Apple gave iOS developers a remarkably simple API for something genuinely complicated happening underneath it. Simple API, complicated reality — that gap doesn’t close just because you can’t see it. Test on hardware that isn’t yours, keep @Generable doing what it’s actually good at, and stop being surprised when “the same model” gives two different answers on two different phones. It was never actually the same computation. It just looked that way from the API surface.

If you’re building anything with Foundation Models beyond a toy, ThinkBud is where I’ve been stress-testing exactly this — on-device inference with zero server round-trip, which means zero place to hide when a device-specific answer goes sideways.

For more on where Apple’s on-device tooling has genuine rough edges, see the Skills API for Foundation Models and the AssetInventory locale bug I ran into a few weeks back. And if you’re relying on AI output anywhere in your workflow, not just in your shipped app, Vibe Coding Native iOS Apps covers the discipline that keeps that reliance honest.

Share this post

Share on X LinkedIn

Comments

Leave a comment

0/1000

N

NativeFirst Team

Editorial

The NativeFirst team — engineers and designers building native Apple apps and writing the courses we wish we had when we started.