The news
GitHub Security Lab said on September 28 that researchers using its open-source AI security agent, the Taskflow Agent, have found and reported 24 vulnerabilities in Android apps, including a navigation app with more than 10 million downloads and the Wikipedia app.
GitHub, owned by Microsoft (MSFT), describes the Taskflow Agent as a framework that lets security researchers automate, package and share the AI prompts and workflows that work for them. Researcher Kevin Stubbings built custom taskflows for Android that split an audit into steps: find an app's entry points, separate the mobile-specific ones, then check each against the vulnerability classes that apply to it. GitHub's post does not name the AI models used. It says a GitHub Copilot license is required, that the prompts use premium model requests and that runs can consume a large number of tokens.
One example involved OsmAnd, a navigation app with more than 10 million downloads on the Play Store. According to Help Net Security, an exported screen in the app accepted instructions that should have stayed internal, so any other app on the phone, even one with no permissions, could silently import malicious settings. GitHub said that let an attacker receive private location data, including the origin and destination of every route a user took.
A second example was an account takeover flaw in the Wikipedia Android app. Help Net Security reported that a hostname check in the app's deep link handler matched only the end of a domain name rather than the full domain. GitHub said an attacker could then obtain the victim's username, a long-lived token and a session token valid across every Wikimedia project.
Other bug types GitHub listed include path traversal, cross-app scripting through WebView, exposed JavaScript bridges, flaws in how apps handle intents from other apps, deep link logic bugs and cookie leakage. The post does not say whether all 24 flaws have been fixed, and it points readers to GitHub's advisories page for disclosures.
Stubbings was frank about the limits. He wrote that large language models are good at finding vulnerabilities but struggle to estimate severity, and that the AI sometimes flagged issues requiring conditions almost impossible to meet in real life. GitHub said each finding should be reviewed by a security researcher who understands mobile apps.
The numbers
- Vulnerabilities found and reported so far
- 24 (GitHub Security Lab)
- OsmAnd downloads on the Play Store
- More than 10 million (Help Net Security)
- Typical taskflow run on a medium-sized repository
- An hour or two (GitHub)
Why CEOs should care
For CISOs and application security leaders, the main point is access. The tooling is open source, runs on a commercial AI subscription and, according to GitHub, can take an hour or two on a medium-sized code repository. That puts structured AI auditing within reach of in-house teams, and of attackers. Companies that ship Android apps should run this kind of review on their own code first, focusing on the areas where GitHub found flaws: exported components, links that open the app, WebView bridges and data passed between apps.
For CFOs, the cost is not only the license. GitHub warns that runs can consume large numbers of tokens, and every finding still needs review by a mobile security specialist, because the models misjudge severity and raise findings that would rarely matter in practice. Budget for triage time as well as compute, or the output becomes an unread backlog.
For product leaders and boards of consumer app companies, the Wikipedia example is the one to study: a single loose domain check led to account takeover. Ask when your mobile apps were last audited for flaws of this kind, and whether your vulnerability disclosure process can handle more incoming reports as outside researchers adopt AI tools.
The bigger picture
GitHub's conclusion is that AI will become an essential tool for open-source maintainers in both development and security. Its method is worth noting: rather than asking a model to find bugs in a whole codebase, the taskflows break the job into small, guided steps. Stubbings wrote that even as new models get better at understanding code, custom prompts let researchers steer them one stage at a time. He also expects false positives to fall as models handle more context and reason better, which would lower the triage burden.
What’s next
GitHub described the 24 flaws as found so far, and its advisories page will show further disclosures as they are published. Watch for fixes from affected app makers, and for other security teams publishing their own taskflows for mobile and web code.
What “Fact-checked” means
Fact-checking means testing a story’s facts against the evidence before it is published. This story went through at least two separate checks before this version was published.
- What we checked
- Its names, figures, dates, job titles, quotes and who said what were checked against the story’s sources, including its main source where it could be opened. The headline was checked for accuracy and overstatement.
- How
- A first check reviewed the whole story. If it passed, a second, skeptical check went back to the sources to look for mistakes in the most important facts. If a check flagged the story, it was edited to fix the problems found, and a separate re-check then reviewed the whole story again.
- Who
- The checks are made by our newsroom, as steps kept separate from the writing, under rules set by our editor, Hussein Mukhtar. A story the checks still flag is held for the editor, who decides whether it is fixed, published or dropped.
- If something is wrong
- “Fact-checked” does not mean error-free. If a material error is found after publication, we correct the story and add a note saying what changed. Report an error
Companies in this story








