Your Crash Rate Doubled Overnight. You Didn't Ship Anything.
There is a specific flavor of dread that comes from opening your crash dashboard and seeing a cliff. Not a slope — a cliff. Crash-free sessions fell off a table at some point overnight, and your last release was eleven days ago.
Your first instinct is to check whether you actually shipped something and forgot. You didn’t. The binary on every one of those phones is byte-for-byte the same binary that was fine yesterday.
The thing that broke was never in your app
Last week on r/swift, someone posted a thread titled “Firebase Analytics suddenly causing production iOS crashes with no app release.” The phrasing is doing a lot of work there, and every word of it is the interesting part. Not “after I updated Firebase.” Not “after I shipped 2.4.0.” With no app release.
I can’t independently verify what happened inside Google’s infrastructure that day, and I’m not going to pretend otherwise. But the shape of the failure is one I’ve seen enough times to describe precisely, because it’s structural. It isn’t a Firebase thing. It’s a consequence of how basically every analytics, experimentation, crash-reporting, and feature-flag SDK on iOS is designed.
Here’s the shape. Most of these SDKs don’t just send data up. They pull configuration down. On launch, or shortly after, the SDK hits its backend and fetches a payload: which events to sample, which experiments this device is in, what the current rate limits are, which endpoints to use, what the schema version is. Then it parses that payload and reconfigures itself.
That payload is code’s evil twin. It isn’t code, so it never went through your review, your CI, your TestFlight build, or App Review. But it absolutely determines control flow inside a library running in your process, with your entitlements, on your main thread.
So when someone on the other end of that connection rolls out a change — a new field, a reordered enum, a different type for a value that used to be a string, a flag that enables a code path that’s been dormant for two years — the parsing or the newly-enabled path can hit an edge case. Force-unwrap a key that isn’t there. Decode a payload shape nobody tested on your SDK version. And your perfectly healthy shipped app takes a SIGABRT in a framework you never wrote.
You didn’t ship a bug. Someone shipped a bug into your app.
Why pinning your dependency versions does nothing here
This is the part that catches people, and it’s worth being blunt about it.
Every piece of standard dependency hygiene we’ve all internalized is about controlling which version of someone’s code lands in your binary. Pin exact versions in Package.swift instead of .upToNextMajor. Commit your Package.resolved. Review the diff when you bump. Don’t let a transitive dependency float.
All good advice. None of it helps you here, and it’s important to understand why: version pinning protects you from changes to the SDK’s code. This is a change to the SDK’s data.
// You did everything right.
dependencies: [
.package(
url: "https://github.com/firebase/firebase-ios-sdk.git",
exact: "11.4.0" // pinned, resolved, committed, reviewed
)
]
That exact: is real protection against a category of problem. It is zero protection against the SDK you pinned asking a server a question and getting a new answer. The binary is frozen. The input isn’t.
Same reason your staged rollout doesn’t save you. Phased release, 1% then 10% then 50% — that’s a mechanism for limiting the blast radius of your binary. When the change isn’t in your binary, every user on every version you’ve ever shipped that still contains that SDK gets it simultaneously. There is no 1% ring. The rollout already happened, and it happened to everybody.
And for the same reason, rolling back doesn’t work the way your reflexes want it to. You can’t un-ship it. The fastest release you’ve ever cut still needs App Review, and the thing causing the crash will keep being served the whole time.
Why it takes so long to even identify
The diagnosis is miserable in a very particular way, and it burns hours.
The stack trace points into a binary you don’t have symbols for. The top frames are inside the SDK, often partially symbolicated, sometimes just addresses and a framework name. Your own code appears somewhere down around frame fifteen, in whatever innocent line called configure() at launch. So the trace tells you where the crash surfaced and almost nothing about why.
Then your instincts actively work against you. You go looking for what changed on your side, because that’s what debugging is — find the delta. There is no delta on your side. You can spend a genuinely embarrassing amount of time diffing a release against itself before the thought “what if it isn’t us” arrives.
It gets worse if the crash is a launch hang rather than an exception. An SDK doing synchronous network work during startup can blow the watchdog timeout and get your app killed with 0x8badf00d — which lands in your dashboard looking like a crash but is really the system losing patience. Those have no useful exception at all, just a stack of threads all waiting.
This is honestly one of the better uses for an AI assistant in a debugging loop — not to find the bug, but to help you read an unfamiliar framework’s stack trace and reason about what a third-party binary is doing at launch. I wrote about that workflow in more detail in debugging with AI. It won’t tell you Google changed a payload. It will stop you from re-reading your own diff for the fifth time.
What actually helps
The honest summary is that you cannot prevent this. You can only shorten it. Four things genuinely do that.
Ship your own kill switch, before you need one. This is the big one. Every third-party SDK in your app should be behind a flag that you control, checked before you initialize it:
// In your own config — a server you own, or a cheap JSON file on your CDN.
if remoteConfig.isEnabled("analytics_sdk") {
FirebaseApp.configure()
}
That’s a few lines, and it converts “wait for App Review” into “flip a value, crashes stop in minutes.” The catch is that it only works if it’s already in the shipped binary. A kill switch you add in response to an incident is not a kill switch, it’s a hotfix with extra steps. Add them while nothing is on fire.
Make sure your own config fetch is the one thing that can’t take the app down with it — wrap it in a timeout, cache the last good value, and default to whatever keeps the app running. The point of the switch is to be more reliable than the thing it’s guarding.
Get optional SDKs off the launch path. Analytics does not need to initialize before your first frame. Neither does most attribution, experimentation, or support tooling. Defer them to after launch, off the main thread where the SDK permits it. An SDK that crashes during a background initialization twenty seconds in is a bug report. The same SDK crashing in didFinishLaunching is an app that cannot be opened — and your users’ only available fix is deleting it. The same reasoning applies to anything you schedule at startup; I went through the launch-path cost of background work in the BGTaskScheduler piece.
Alert on crash-free rate continuously, not around releases. Most teams watch stability like a hawk for forty-eight hours after shipping and then stop looking. That monitoring model has an exact blind spot for this failure, which arrives on a random Tuesday with no release attached. Set a threshold alert that fires regardless of whether you deployed.
Count your SDKs, and know what each one phones home to. Not as a purity exercise — as a blast-radius inventory. Every SDK that fetches remote configuration is an additional party who can change your app’s behavior without asking you. Four of those is a different risk profile than fifteen. It’s the same trust-boundary question as validating purchases on a server you control rather than trusting the client, or the API keys people keep bundling into their binaries: the question is never “is this library any good,” it’s “what can this library do to me, and who else gets to decide.”
When I was deciding what went into Invoize, this was the whole argument for keeping the dependency list almost offensively short. Not minimalism as aesthetics. Fewer third parties who can change what my shipped app does on a Tuesday while I’m asleep.
Your binary is not your app
That’s the mental model shift worth keeping.
What runs on your users’ phones is your compiled binary plus every remote payload every SDK in it fetches at runtime — plus whatever those payloads switch on. You own the first part completely. You rent the rest, from vendors with their own release schedules, their own incident response, and no obligation to tell you when they’re rolling something out.
Most of the time that trade is worth it. Writing your own analytics pipeline to avoid this is not the lesson. But it’s worth knowing which parts of your app you actually control, because when the cliff shows up in your dashboard at 7 AM, the useful first question isn’t “what did I break.”
It’s “what did somebody change.”
Share this post
Comments
Leave a comment
NativeFirst Team
EditorialThe NativeFirst team — engineers and designers building native Apple apps and writing the courses we wish we had when we started.