The short version: Researchers from Wake Forest University analyzed 444 free iOS apps with large language model features and found that 282 of them — 64% — exposed exploitable AI provider credentials or unauthenticated backend access in plain network traffic, with no jailbreak or decompilation required. The leaks fell into three patterns: hardcoded plaintext keys, open backend relays that need no credential at all, and "temporary" access tokens that behaved just like raw keys once captured. Of the 282 apps, 146 were fully exploitable, letting an attacker run inference on the developer's bill or pull out proprietary system prompts. Three months after responsible disclosure, only 28% of affected developers had actually fixed the problem.
What the Wake Forest researchers found
The study, titled "Mind your key: An Empirical Study of LLM API Credential Leakage in iOS Apps," set out to measure how many AI-assistant apps on the App Store properly protect the credentials they use to call OpenAI, Google Gemini, Anthropic, and similar providers. The team started from more than 38,000 App Store listings and narrowed that pool down to 444 free apps confirmed to have working LLM features, which they then tested on physical devices without jailbreaking anything.
Testing showed that 64% of those apps exposed credentials or access paths that an outsider could actually exploit. The vulnerable apps weren't clustered in one niche — they spanned thirteen categories, from productivity and education to lifestyle and health and fitness. This also wasn't limited to obscure, low-traffic apps: roughly 15% of the vulnerable apps had over a thousand user ratings, and the single most popular affected app had collected more than 2.3 million ratings.
How you pull an AI key out of an app without touching the binary
iOS complicates static analysis in a way Android doesn't. App Store binaries are wrapped in Apple's FairPlay DRM, so decompiling them the way researchers routinely decompile Android APKs isn't practical without a jailbroken device. So the Wake Forest team didn't bother trying. They built a tool called LLMKeyLens and simply watched what each app said over the network.
Each app was installed on a real device and run through a man-in-the-middle proxy with a custom root certificate, decrypting its HTTPS traffic while a tester triggered the app's AI features with controlled prompts. LLMKeyLens then fingerprinted which LLM provider each request was headed to, pulled out anything that looked like a credential, and tried it against the live service to confirm it actually worked.
The choice of method matters as much as the headline number. Whatever the researchers found was, by definition, something that had already left the device in the clear — it didn't matter how the key was stored, obfuscated, or constructed inside the app itself. A key hidden behind string-splitting tricks or assembled at runtime still has to leave the device in an HTTP header to be useful, and that's exactly where this study looked.
1 · Candidate discovery
Researchers narrowed an initial pool of 38,520 App Store listings down to 444 free apps with confirmed, exercisable LLM functionality.
2 · Traffic capture
Each app ran on a physical device behind a MITM proxy while testers triggered its AI features, capturing every outbound request the app made.
3 · Credential validation
Suspected credentials were matched against known provider patterns and then validated with benign, real requests to confirm they still granted live access to the AI service.
4 · Disclosure and recheck
The team notified affected developers and, checking back three months later, found only 28% had actually fixed the problem.
Three ways the keys got out
The 282 vulnerable apps split into three distinct failure modes, and the split is instructive: it shows that the "obvious" fix — stop hardcoding the raw key, hand out a token instead — wasn't sufficient on its own.
- Plaintext keys (54 apps). The provider's static API key was visible in the open, readable from a single captured request.
- No key needed (92 apps). These apps routed requests through a backend that answered any caller with no check on identity — effectively an open relay sitting on top of a paid AI account.
- Replayable tokens (136 apps, the largest group). Instead of a raw provider key, the app handed back a session or access token — the approach teams reach for because it sounds safer. In practice the tokens still crossed the wire unprotected and worked exactly like a static key once intercepted: capture once, replay indefinitely.
All 282 of the flagged apps were confirmed to be actively exploitable through credential replay at the time of testing, not just theoretically vulnerable. Of those, 146 cases — just over half — involved either a plaintext key or a fully unauthenticated backend, meaning the attack required nothing beyond capturing and replaying a single request.
What an attacker actually does with a stolen AI key
A leaked LLM API key isn't a curiosity — it's a line of credit. These credentials function like passwords that let an app talk to an AI model, and once copied, an outsider can reuse them while the original developer keeps paying for the resulting usage. Whoever captures the key or token can send model requests under the developer's account, and the developer is the one who gets billed.
Three consequences follow directly from the way these apps were built:
- Unmetered usage on someone else's bill. Any rate limit the app's UI enforces — a daily chat quota, a usage counter — is cosmetic once an attacker calls the provider directly with the harvested credential. There's no app-side throttle on a request made outside the app.
- Business logic walks out the door. Beyond running up usage, attackers who reach these backends can disrupt the app or extract the hidden system prompts that encode its behavior. System prompts routinely contain prompt engineering and guardrail logic a team spent real time building.
- Open relays implicate more than the developer. Where the leak is an unauthenticated backend rather than a raw key, that backend is often the same server that proxies every user's requests. A single discovered endpoint can become a standing, anonymous path into a shared AI backend — not just free usage, but potentially a window onto how that backend handles other users' prompts and responses, depending on how it's built. The study didn't set out to measure cross-user data exposure specifically, but the architecture it documents makes that risk structurally possible wherever one proxy serves many users behind a single shared credential.
Why this keeps happening
LLM features got bolted onto mobile apps fast, and the quickest way to call OpenAI or Anthropic from client code is to put the key where the code can reach it. That's the same shortcut that produces hardcoded secrets in build pipelines and server environments — see our piece on credential attacks no checklist can stop — except a mobile binary ships directly into the hands of anyone who wants to inspect it. There's no perimeter between "inside the app" and "in front of the attacker"; they're the same place.
The "replayable token" category is the more interesting failure, because it shows teams that tried to do better. Handing the client a token instead of a raw provider key is a reasonable instinct. It only holds up if the token is short-lived, scoped to one user and one action, and bound to something an attacker can't simply replay from a packet capture. A token that's valid indefinitely and works for anyone who holds it is a key with extra steps.
This is the structural gap KnoxCall's AI gateway is built around: the mobile app, the CI runner, or the agent process should never hold the actual OpenAI or Anthropic credential in the first place. Instead of embedding a key and hoping it's never captured, requests go to an egress proxy that attaches the real provider credential at the wire, after verifying the caller, so the app itself only ever holds a scoped, revocable handle to that proxy. Intercept that traffic and you get a narrow, short-lived permission — not the vendor's master credential for the account.
To be clear about what that does and doesn't solve: wire-level injection doesn't make interception impossible, and a badly scoped proxy token can still be abused the same way this study's replayable tokens were. The real gain is in blast radius and revocability — a leaked per-call token can be invalidated or rescoped without rotating the underlying provider key everywhere it's used, and the provider credential itself never has to live anywhere an attacker can read it from a mobile binary or a traffic capture. That's a narrower, more honest claim than "this would have stopped the leak" — it's a claim about what's exposed when a leak happens anyway.
What to do if your app calls an LLM provider directly
- Get the provider key off the client, full stop. If your app sends requests straight to an LLM provider, whatever credential rides along in that request is observable to anyone capturing traffic, no matter how it's obfuscated inside the binary.
- Route LLM calls through a server or gateway you control. The client authenticates to your own backend; your backend attaches the real provider credential, not the device.
- Make client-facing tokens short-lived and scoped. A token that works forever, for any caller, is functionally a static key — see our guide on common API key security mistakes.
- Rate-limit and monitor per user, not just per app. Usage anomalies tied to a specific account are the first real signal of a stolen credential.
- Rotate and verify, not just announce. In this study, only 28% of affected developers had properly fixed the issue 90 days after being notified — a patch note is not the same as a verified fix.
None of this is iOS-specific. Android apps carry the same incentive to hardcode keys, and earlier research on Android LLM apps found the same core problem at scale, including hardcoded, provider-specific credentials shipped directly inside app binaries. The channel is the mobile network request; the platform is close to incidental.
The broader pattern
This is not the first time researchers have gone looking for hardcoded credentials in mobile apps and found them at scale, and it won't be the last. What makes this study worth reading closely is the method: no binary teardown required, just a proxy and some patience. It's notable how little effort the discovery took — the researchers' tool simply watched each app's traffic and extracted the credentials as they passed by. If a credential is observable from outside the app, obfuscating it inside the app was never the control that mattered. See our broader security overview for how egress-level credential handling fits into API security more generally, and our gateway comparison if you're weighing how to make this change without rebuilding your entire backend.