A 16-Year-Old SQLite Bug Just Surfaced. Your SwiftData App Is Sitting on the Same Engine.

NativeFirst Team 9 min read
An open drawer of a vintage card catalog cabinet, packed with densely filed index cards

There’s a bit in The Big Short where Steve Carell’s character keeps flying around asking people why the mortgage bonds are rated AAA, and everyone gives him a slightly different non-answer, and you can watch the exact moment he realizes nobody actually checked. That’s roughly the feeling of reading Tailscale’s post-mortem yesterday and then going to check which SQLite your iPhone app is running on.

Tailscale published the story of a 16-year-old SQLite bug that ate six months of their uptime. It hit the top of Hacker News with 1,000+ points. The short version: 19 separate database corruptions over six months, months of forensics, and at the end of it a data race in SQLite’s write-ahead log that had been sitting there since 2010.

SQLite is the most-tested piece of software most of us will ever depend on. It’s also the engine underneath SwiftData and Core Data. So I spent ten minutes checking what my own machine ships. I’ll get to what I found.


What the bug actually is

SQLite’s WAL — write-ahead log — is the thing that lets one process write while others read. Instead of editing the main database file in place, writes append to a -wal sidecar file, and periodically a checkpoint folds those changes back into the main file and resets the log.

The bug lives in that reset. From the SQLite team’s own note when they shipped the fix:

“The bug is a data race with tight timing constraints. It is unlikely to occur in common use.”

They went further and admitted they had never reproduced it organically — they had to add deliberate testing logic to SQLite that forces the exact circumstances, just to confirm the fix worked. That is a remarkable sentence to read about a database with a test suite famous enough to have its own marketing page.

Fixed in SQLite 3.51.3. Introduced in 2010. Sixteen years of production traffic in between.

The reason nobody found it earlier is the reason it’s interesting: it needs a writer and a checkpoint racing each other with precise timing, on a database getting hammered continuously. Tailscale runs a single-writer Go process per shard against SQLite — textbook usage, exactly how the docs say to do it — and they still got bitten 19 times, because at their volume “unlikely” happens on a schedule.

There’s a great companion piece from Antithesis, who took the still-buggy 3.51.2, pointed their deterministic simulator at it with a completely generic workload — concurrent writes and checkpoints, nothing exotic — and caught it in 15 minutes. Same bug that took Tailscale six months of production forensics. The workload wasn’t clever. The assertions were the boring ones: no lost committed writes, database is not corrupt. Their line about it is the one I keep thinking about: the simplest workloads find the hardest bugs.


Now the part that concerns us

SwiftData is not a database. Core Data is not a database. They’re object graph layers, and underneath both of them, for the default store type, is SQLite. Apple’s persistent store has used WAL journaling by default for years — that’s what those -wal and -shm files next to your .sqlite in the app container are.

So: which SQLite does iOS ship?

I checked the two I have on this machine:

$ sqlite3 --version
3.51.0 2025-06-12 13:14:41 f0ca7bba1c5e232e5d279fad6338121ab55af0c8c68c84cdfb18ba5114dcaapl (64-bit)

$ strings "/Library/Developer/CoreSimulator/Volumes/iOS_23C54/.../RuntimeRoot/usr/lib/libsqlite3.dylib" | grep -E '^3\.[0-9]+\.[0-9]+$'
3.51.0

macOS 26: 3.51.0. The iOS 26.2 simulator runtime: 3.51.0. The fix landed in 3.51.3.

Before you go rewrite your persistence layer, read the next section, because I want to be honest about what that does and doesn’t prove.


The honest caveats, because this is where blog posts usually cheat

Apple ships a fork. Look at that version string again: it ends in dcaapl. That apl suffix is Apple’s build of SQLite, not the upstream one. Apple routinely backports security and correctness patches without moving the version number. So “3.51.0” tells you the baseline, not the patch level. It is genuinely possible the WAL-Reset fix is already cherry-picked into that binary and the string just doesn’t say so. I can’t see inside it, and neither can you.

Your app is not Tailscale. This bug needs a writer and a checkpointer racing under sustained concurrent load. A typical iOS app has one process, modest write volume, and long idle periods where the user is, you know, not using the app. The odds of hitting a data race that SQLite’s own developers couldn’t trigger on purpose are somewhere between “low” and “buy a lottery ticket instead.”

But “low” isn’t “zero,” and some apps aren’t typical. If you’re running a background sync that writes while the foreground writes, or you’ve got an app group with multiple processes touching the same store, or you’re doing a large import while the UI keeps querying — congratulations, you’ve built a smaller version of the workload that Antithesis used to reproduce this in 15 minutes.

So no, I’m not telling you the sky is falling. I’m telling you that “the storage layer is somebody else’s solved problem” is an assumption, and this week is a decent reminder to know what you’re actually standing on.


What’s actually worth doing

Run an integrity check somewhere. Not on every launch — it’s not free on a large store — but a debug-menu button or a once-per-release diagnostic costs you nothing and would have told Tailscale on day one:

import SQLite3

func integrityCheck(at path: String) -> String {
    var db: OpaquePointer?
    guard sqlite3_open_v2(path, &db, SQLITE_OPEN_READONLY, nil) == SQLITE_OK else {
        return "could not open"
    }
    defer { sqlite3_close(db) }

    var stmt: OpaquePointer?
    guard sqlite3_prepare_v2(db, "PRAGMA integrity_check;", -1, &stmt, nil) == SQLITE_OK else {
        return "could not prepare"
    }
    defer { sqlite3_finalize(stmt) }

    guard sqlite3_step(stmt) == SQLITE_ROW,
          let cString = sqlite3_column_text(stmt, 0) else { return "no result" }
    return String(cString: cString)   // "ok" when healthy
}

Point it at your store URL. Healthy databases return the string ok. Anything else is the conversation you want to have before your users do.

Know where your store actually lives. A surprising number of people shipping SwiftData apps have never looked at the files. modelContainer.configurations.first?.url gives you the path; the -wal and -shm siblings sitting next to it are the write-ahead log this whole story is about.

Have a corruption path that isn’t a crash. Tailscale’s recovery took over an hour in the early incidents. Yours shouldn’t — for most apps, “detect it, wipe the cache store, re-sync from the server” is a perfectly good answer, and it’s a much better answer than an infinite launch-crash loop that gets you one-star reviews you can’t reply to.

If you’ve been putting off the “what if the local store is garbage” branch because it felt paranoid: it’s not paranoid, it’s just rare, and rare things happen to apps with users.


The part that isn’t really about SQLite

What sticks with me isn’t the bug. It’s the shape of the story.

Tailscale did everything right. Boring technology, single-writer design, exactly the usage pattern the documentation recommends, backups every few minutes since 2023. And they still spent six months writing a custom transaction logging pipeline and shimming a debugging tool into SQLite’s virtual filesystem layer just to see what was happening. Their own line: nobody wanted them to spend six months looking for bugs in SQLite.

Meanwhile the reproduction, once someone pointed the right kind of tooling at it, took a quarter of an hour.

That gap — six months of production forensics versus 15 minutes of deterministic simulation — is the whole argument for investing in the boring parts of your toolchain before you need them. It’s the same thing I was circling last week when Hacker News argued about whether code was ever the hard part: writing the code was never the bottleneck. Finding out why the code you already wrote is behaving impossibly — that’s the job.

And it’s the same lesson as the strict concurrency migration, where the compiler pointed at a race condition I’d been shipping for months without knowing. Races don’t announce themselves. They wait for load.


If you want the “SwiftData is a layer over SQLite and that has consequences” version of this argument at more length, I migrated two apps to SwiftData and moved half of it back, and SwiftData two years later covers what’s solid versus what still bites. If the part you want to get better at is the hunting — reading a stack trace, forming a hypothesis, actually narrowing it down instead of changing things until the red goes away — the Debugging with AI lesson in the course is about exactly that workflow.

Go check your integrity_check. It takes a minute, and it’s a better use of today than arguing about Grok on Hacker News.

Share this post

Share on X LinkedIn

Comments

Leave a comment

0/1000

N

NativeFirst Team

Editorial

The NativeFirst team — engineers and designers building native Apple apps and writing the courses we wish we had when we started.