Skip to content
Security & Trust

GitHub's AI agent found 24 Android flaws, and needed experts to judge them

GitHub Security Lab's results suggest the bottleneck in AI-assisted audits is deciding which findings matter.

W
WebPulse Newsroom
AI-assisted · 4 min read
Share on X LinkedIn
GitHub's AI agent found 24 Android flaws, and needed experts to judge them

AI-generated image for WebPulse. About our images

Key finding

Android vulnerabilities found and reported: 24 (Source: GitHub Security Lab blog (September 28, 2026))

When a tool can search code for flaws in an hour or two, finding them stops being the hard part. The hard part becomes deciding which findings are real, which are serious and which can be ignored. That is the lesson in a September 28 write-up from GitHub Security Lab. The difficulty of security research did not shrink. It relocated, from discovery to judgment.

What GitHub reported

GitHub's Security Lab built the Taskflow Agent, an open source framework for packaging AI prompts and workflows into repeatable audit steps. A researcher there adapted it for Android applications and reports 24 vulnerabilities found and reported so far. The write-up says the team found a handful of critical ones, many simple ones such as path traversal, and that the types of flaws were where an experienced researcher would expect them. It also describes Android app security as fairly robust overall.

The adaptation was human work. The researcher added a step that separates mobile entry points, the places attacker-controlled data can enter, from non-mobile ones. The researcher also edited a classification step to name Android-specific vulnerability classes the model must check, because mobile flaws are less widely known and models are non-deterministic. The AI supplied scale and creativity. A specialist supplied the map.

24
Android vulnerabilities found and reported
Source: GitHub Security Lab blog (September 28, 2026)

Two cases that show the stakes

The first involves OsmAnd, a navigation app whose Android version has over 10 million downloads, according to GitHub. An exported screen in the app accepted settings instructions that the developers expected to arrive only from a trusted internal channel. GitHub reports that any app, even one with no permissions, could use this to change OsmAnd's settings without the user noticing. That included redirecting map tile requests to an attacker's server, which would reveal the coordinates of tiles the user had loaded. The same flaw could expose the origin and destination of routes.

The second involves the Wikipedia Android app. GitHub describes a logic bug in how the app parses hostnames in wikipedia:// links, so a page on a lookalike domain ending in wikipedia.org could load as if it were genuine. GitHub says the app then sent the user's cookies, giving an attacker the username, a long-lived token and a session token valid across Wikimedia projects. Chained with a second instance of the same pattern, this became an account takeover. GitHub calls both examples disclosed.

Over 10 million
Android downloads of OsmAnd, per GitHub
Source: GitHub Security Lab blog (September 28, 2026)

Where the work moved

GitHub is candid about the limits. The model returned low-severity bugs even when told not to, and some findings depended on situations the author considered very unlikely to arise in practice. Severity was often estimated incorrectly, because mitigating factors are hard for a model to see. GitHub's example is a path traversal restricted to external storage, which is low severity in practice.

Asking the model to build a proof of concept helped, but required extra runs, and the model could still be wrong. GitHub describes a case where internal storage data takes priority over attacker-writable external data, which means no vulnerability exists at all. GitHub's takeaway is that no finding should skip a human check by someone versed in mobile applications.

An analogy helps here. A smoke detector sensitive enough to catch every fire will also sound for toast. The detector is not the expensive part of the system. The person who walks to the kitchen and decides whether to call the fire brigade is.

1-2 hours
Typical run time on a medium-sized repository
Source: GitHub Security Lab blog (September 28, 2026)

The decision this forces

For an executive, the implication is about budgeting, and it is this publication's interpretation rather than GitHub's claim. If discovery takes hours, the cost of an AI-assisted audit shifts toward reviewers who can separate a real account takeover from a noisy finding. GitHub also notes that the taskflows require a Copilot license, use premium model requests, and can consume a large number of tokens. Spending grows in two places: compute and expert attention.

GitHub itself is bullish on AI-assisted research as a route to safer open source projects. That is the publisher's stated belief, and this one write-up does not establish it. What it does document is a workable method, with the caveats stated openly.

Questions for your security team

Who reviews AI-generated findings, and do they know the platform in question? GitHub's mobile reviewer needed Android expertise, and a general reviewer may not spot a misjudged severity.

Do your mobile applications get audited at all, and by whom? The OsmAnd and Wikipedia cases both stem from how an app trusts input arriving through its own entry points, which is a design question as much as a coding one.

When a tool reports a finding, is a proof of concept required before it enters the remediation queue? GitHub's experience is that this step exposes mitigating factors, at the cost of extra runs.

If a run takes an hour or two, the scarcer resource is the person who can tell which findings matter. That is where the budget conversation belongs.

Produced by the WebPulse Newsroom with AI assistance from the original reporting credited below, and checked against that source by our editorial review. How we use AI.
Original reporting: GitHub.

Share this insight