Trending: On-device modelsSearch
iHeartGeek
iTECH

GitHub's AI security agent found 24 Android bugs

GitHub's Security Lab has open-sourced the taskflow agent behind 24 Android vulnerability reports, including a chained Wikipedia account takeover that turned on a single bad domain check.

Abstract illustration of glowing code streaming past a bright scanning beam across a dark circuit pattern, representing an automated audit of application source code.

GitHub's Security Lab has published an open-source agent that audited Android applications and came back with 24 security holes, including a chain that lets an attacker take over a Wikipedia account. The agent, and the prompt files that steer it, are available for anyone to run against their own code.

The tool is the GitHub Security Lab Taskflow Agent, built so that security researchers can package and share the AI prompts and workflows they already find useful instead of rebuilding them for every engagement. The Android work is a set of auditing taskflows written on top of it, designed to make a language model look at mobile code in the order a human auditor would.

How the Android taskflows work

Two custom taskflows do the heavy lifting. The first sorts an application's entry points into mobile and non-mobile buckets, so a repository holding a phone app alongside a web service does not confuse the model about where the attack surface actually is. The second hands the model a list of vulnerability classes to work through for each entry point and component, which matters for Android because the interesting bugs tend to be specific: if an activity is reachable through an intent, the taskflow asks about intent-shaped problems such as confused deputy flaws and insecure broadcasts.

Running it is deliberately unglamorous. A researcher opens a codespace on the seclab-taskflows repository, runs a shell script against the target organisation and repository, and waits. A medium-sized codebase takes an hour or two, and the results land in a SQLite database with a column flagging which findings the agent believes are real. A GitHub Copilot licence is required and the prompts burn premium model requests, which the team is upfront about.

An exported activity, and a Wikipedia takeover

The most serious example came from OsmAnd, the open-source navigation app with more than ten million Android downloads. OsmAnd exports an activity called MapActivity, which handles settings files and deep links. Because the activity is exported, any other app on the device can send it an intent with any extras it likes, and the settings importer trusts them. A malicious app can therefore push a settings file into OsmAnd silently, replace the map tile URL template with one pointing at its own server, and then read the tile coordinates the app requests. Tile coordinates are coordinates: the attacker learns where the device has been, and the same technique exposes the origin and destination of every route the user plans. Three vulnerabilities were reported in OsmAnd from this work.

The second example is worse for anyone with a Wikipedia account. The Wikipedia Android app registers a wikipedia:// deep link, and its hostname check only asks whether the incoming authority string ends with the Wikipedia domain. A hostname such as evil-wikipedia.org passes that test, so a link on a web page can open the app, show an attacker-controlled page inside a WebView that looks like Wikipedia, and run attacker JavaScript in it. A second copy of the same ends-with logic in the app's cookie manager hands over the cookies for wikipedia.org. Chained together, the two bugs leak the long-lived session token that works across every Wikimedia project.

What the agent still gets wrong

GitHub is careful not to oversell the results, and the caveats are the most useful part of the write-up. The agent is good at finding vulnerabilities and poor at judging how much they matter, frequently reporting low-severity issues, returning findings that need very specific conditions to exploit, and misreading mitigating factors that reduce real-world impact. The team's fix is more prompting: asking the model to build a proof of concept forces it to try to exploit what it found, which surfaces the false positives. Every finding still needs a human who knows mobile code to review it.

Our opinion

The headline is 24 vulnerabilities and the actual news is the scaffolding. Nothing here depends on a smarter model: the gains come from splitting an audit into steps, telling the model what an Android entry point is, and forcing it to work through vulnerability classes it would otherwise skip. That is an admission that generic agentic coding has hit a wall on security work, and it is a more honest position than the vendors claiming a model can audit a codebase unaided. It is also telling that the two showcase bugs are not exotic memory corruption but trust mistakes that a careful human reviewer should already be catching: an exported activity that trusts extras, and an ends-with domain check applied to a hostname and then again to a cookie. Both are boring, both are everywhere, and both are exactly the kind of thing that never gets fixed because nobody has time to look. If an agent can make that look worth the tokens, it has earned its place in a review pipeline, provided someone senior still signs off on the severity. The Wikipedia chain in particular should be read as a warning about deep links generally rather than a story about one app.