# A Firebase Payload Crashed Released iOS Apps Without an Update

**Summary:** An incorrectly formatted Firebase Analytics payload crashed already-released iOS apps, then cached responses lingered for up to four hours. I would make nonessential telemetry unable to block launch.

- Canonical: https://markhuang.ai/news/firebase-payload-crashed-released-ios-apps
- Language: en
- Author: [Mark Huang](https://markhuang.ai/about)
- Published: 2026-09-29
- Section: News
- Tags: Firebase, iOS, Mobile SDKs, App Reliability, Google Analytics
- Source: [Gergely Orosz on X](https://twitter.com/GergelyOrosz/status/2104825886922911981)
- License: https://creativecommons.org/licenses/by-nc/4.0/

---

![A malformed data stream fans out from a cloud service and crashes a field of smartphones](https://cdn.markhuang.ai/news/firebase-payload-crashed-released-ios-apps/hero.webp)

*The dangerous part of a mobile SDK is not only the code shipped in the app. It is the live behavior the vendor can still change after release.*

On September 29, 2026, Gergely Orosz [called attention to a Firebase failure](https://twitter.com/GergelyOrosz/status/2104825886922911981) that was crashing iOS apps using its Analytics SDK. The linked [Firebase issue](https://github.com/firebase/firebase-ios-sdk/issues/16728) is more precise: one developer saw 56 crashes affecting 56 users in the first 18 minutes, across four already-released builds, without shipping an app update.

What changed was a server response. Google later said an incorrectly formatted payload caused apps to crash on launch. My read is that Firebase Analytics was not behaving like inert library code here. It was a live production dependency with a path from Google's servers to the app process, and a bad response could terminate that process.

That distinction matters more than the SDK version in the issue report. Teams usually treat analytics as something that observes the product. In this incident, the observer could take the product down. I would now review every remotely configured, nonessential SDK with one blunt question: can it stop the app from opening?

## Two hours on the server, four more in the cache

The official timeline says the incident began on September 28 at 17:41 PDT. Google fully rolled out a fix at 19:52 PDT, 131 minutes later. No SDK update was required. Because affected responses could remain cached, Google warned that some app instances might continue crashing for up to four more hours, until 23:52 PDT.

The original report gives the clearest clue that this was not an ordinary release regression. All four affected builds began failing at the same time. In the 42 crashes the reporter inspected, each happened immediately after the SDK received an HTTP 200 response from the `sdk-exp` endpoint. The app then failed while the Analytics experiment worker was processing that response. Google confirmed the broader cause in its pinned resolution: an incorrectly formatted payload received by the SDK.

Orosz described the blast radius as all iOS apps using the SDK and all sessions. The evidence I inspected does not establish that universal scope, so I would not repeat it as a measured total. Public developer reports do show several teams seeing abrupt spikes. In one [iOSProgramming discussion](https://www.reddit.com/r/iOSProgramming/comments/1wt0t3z/firebase_analytics_suddenly_causing_production/), developers reported that crash rates fell after the server fix but lingered while cached responses expired.

## The shipped binary was still changing

No one remotely rewrote those four app builds. The behavior changed because code already inside them accepted fresh input from a service. That makes the usual distinction between a pinned dependency and a hosted dependency less comforting. The package version was pinned, but part of its runtime behavior was not.

This is the product consequence I care about. A team can review an SDK update, test it, stage a rollout, and still inherit a new failure path later through configuration, experiments, flags, models, or data returned by the vendor. Remote control is often the feature. It becomes a liability when malformed input can escape its subsystem and crash the host app.

> **Info:**
>
> My SDK review rule: if a vendor can change behavior after I ship, I treat that integration as a production service, not merely a package. It needs input validation, failure isolation, monitoring, and a way to turn it off before it can block launch.

I made a similar argument after [a Claude service incident](/news/claude-outage-dependency-test), but this case is harsher. A hosted coding tool can disappear and slow work. A mobile analytics SDK can fail inside software that customers already downloaded.

## A kill switch only helps if launch can reach it

The practical suggestion in the public discussion was to put analytics initialization behind a server-controlled switch. I like the direction, with one condition: the app must read that switch before the risky SDK path runs, and the switch cannot depend on the failing vendor.

Firebase's own [iOS documentation](https://firebase.google.com/docs/analytics/ios/configure-data-collection) says Analytics collection is enabled by default. Developers can set `FIREBASE_ANALYTICS_COLLECTION_ENABLED` to `NO` in the app's `Info.plist`, then enable collection later with `setAnalyticsCollectionEnabled`. Google also documents a permanent deactivation setting. Those controls exist, but this incident is a reason to test their timing rather than assume a runtime call will always arrive before an automatic startup path.

If analytics is not required to render the first screen, I would start it disabled and enable it only after the app has loaded its own small, independently hosted policy. The policy fetch must fail safely. A timeout, bad response, or unavailable control service should leave analytics off and let the product run.

That design has a cost. It can lose early-session events, adds another path to test, and may complicate attribution. I would accept that cost for telemetry. Losing a few events is a better failure mode than losing the session itself.

## Google fixed it quickly, but the contract still failed

The strongest defense of Firebase here is fair: Google identified the server-side cause, rolled out a fix in a little over two hours, and spared developers from rushing an SDK release through review. Any large service can have a bad deployment. The server fix was exactly the recovery mechanism customers needed.

But fast recovery does not answer why a malformed experiment payload could become a fatal exception in the host app. Network responses are untrusted input, even when they come from the SDK vendor's own endpoint. The client contract should reject the bad record, retain the last known good state if that is safe, record a diagnostic, and continue without analytics. A nonessential subsystem should not get a vote on whether the app launches.

Communication also needs a cleaner home. The [Firebase status dashboard](https://status.firebase.google.com/) currently directs Google Analytics incidents to a separate Ads Status Dashboard, while the concrete explanation for this event appeared in a GitHub issue. Developers found the answer, but incident discovery should not require following a social post into a repository thread.

## What I would change before the next payload

I would inventory every SDK that fetches remote behavior during startup and record who controls that response, when initialization happens, and what the app does with malformed data. Then I would force those endpoints to time out, return empty objects, send wrong types, and replay stale payloads. The pass condition is simple: the product opens and the optional feature stays off.

I would also monitor SDK failures separately from app releases. Four old builds failing at once was the clue in this case. When crash volume moves without a deployment, the first triage view should include remote configuration, third-party responses, and cached state.

Firebase recovered this incident from the server. App teams mostly had to wait for the fix and the cache window to pass. That is the part I would not leave unchanged. Analytics can help me understand whether a product works. It should never be able to decide that the product does not open.
