SAN FRANCISCO — A drone powered by AI-written software stalks its human target around the house using facial recognition, gliding through doorways and navigating around lamps. It isn’t a scene from science fiction. The cheap consumer drone, flying hands- and (mostly) human-free, is being directed by programs built with widely available AI models from Anthropic, OpenAI and other companies.
The demonstration is part of an evaluation called Drone-Bench, designed by a team of San Francisco researchers at Andon Labs with input from researchers at Anthropic.
The AI systems were prompted (by humans) with simple instructions to fly the drone and track a person based on a picture downloaded from the internet.
Lukas Petersson, one of the co-founders of Andon Labs, said he and his team wanted to highlight just how quickly today’s AI systems are advancing — showing they are capable not only of drafting emails or providing recipe tips but also of shattering the divide between the virtual and physical worlds.
Many people conceived of AI “as a chatbot, and then it became an agent that could do things digitally on your computer,” Petersson said from Andon Labs’ headquarters against a backdrop of spare robot parts and other AI experiments. “We also want to show that, yes, AI can now go into the real world.”
The zeal to translate AI’s growing capabilities into visceral and physical displays is at the heart of Andon’s work. Among the group’s current projects: an AI-run boutique store in San Francisco, an AI-managed café in Stockholm and AI-powered vending machines.
Drones seemed like an obvious next step to test AI’s real-world abilities. Drones have recently become ubiquitous on the battlefield and in police departments, and consumer and government drone operators have long incorporated facial recognition technology.
The FBI recently solicited new proposals for facial-recognition-equipped drones, while Iran-linked hackers claimed in June that they had infiltrated FBI drone networks and identified their facial-recognition schemes. Fears of swarms of cheap, small drones tracking political or civilian targets have also been a recurring theme in science fiction for years.
Drone-Bench — the “Bench” stands for “benchmark” — reveals how those sci-fi conceptions could take off into reality, with today’s consumer-facing AI systems often outperforming human software engineers’ ability to create software that can help a drone fly and track a target.
The evaluation measures AI systems’ ability to master five capabilities crucial to controlling a drone: reconstructing a physical space into a three-dimensional map, localizing where a drone is within that map, navigating throughout the physical space, detecting a person based on a downloaded photograph and tracking that person throughout the space.
The Andon Labs team compared AI systems’ performance with a human engineer’s effort to determine how AI systems stacked up against human abilities. The evaluation did not examine some logistical steps key to flying a drone in the real world, like measuring how often the AI system correctly connected to the drone’s onboard Wi-Fi system to control it or whether adverse weather or crowded skies affected the drone’s performance.
The Andon Labs researchers tested a range of popular AI systems that were released from May 2024 until now, starting with OpenAI’s GPT-4o, continuing through Google’s Gemini 2.5 from June 2025 until Anthropic’s Fable 5 (June 2026) and OpenAI’s GPT-5.6 Sol (July 2026).
The team found the systems have become steadily more capable over time. While GPT-4o matched the standard human performance 14% of the time, on average across all five of the tasks, Fable 5 mirrored or bested human performance 84% of the time.
The researchers also found that the latest AI systems routinely beat human performance on several of Drone-Bench’s five components. Anthropic’s Opus 4.8 and Fable 5, along with OpenAI’s GPT 5.6 Sol, beat the human comparison score on both the detection and tracking steps — a feat the team said was “a bit eerie to watch.”
“I don’t think we could have done this test a year ago and gotten any interesting results,” said Axel Backlund, an Andon Labs co-founder. “Now we can actually show the trajectory of increasing capabilities.”

Despite the AI systems’ ability to control the drones, the Andon team is clear that the human part of the evaluation is not meant to be a flawless, peer-reviewed scientific experiment but instead a good-faith effort to display how a reasonably skilled and determined person could accomplish the same set of tasks.
The team also acknowledges that the current iteration of Drone-Bench does not entirely represent what a true drone novice could achieve with AI’s help or what an AI system could do entirely on its own. For each of the five evaluation stages, the researchers provided the AI systems with working solutions from the previous stage, since errors in earlier stages would have compounded into high failure rates.
Given the systems struggled most with the first stage, in which they were tasked with constructing a 3D model of an office based on a video and turning that model into a two-dimensional map of objects to avoid, the researchers note that no tested AI system completed the entire end-to-end process autonomously — though they expect AI systems will clear that hurdle quickly as they become more capable. As it is, AI systems have beaten human performance on four of the five evaluation components — just not the initial “reconstruct” phase.
The Andon team says the real-world impact of AI-controlled drones will depend both on the ability of AI systems to fully code end-to-end drone programs and the likelihood that AI systems have the desire or intention to carry out malicious activities, whether because of cyberattacks or bad human or AI actors.
The researchers also discovered that the AI systems being tested often resorted to cheating or manipulation to achieve the evaluation’s goal, shortcutting their way to correct outcomes. Callum Sharrock, a technical researcher at Andon Labs, discovered that the AI system being tested had figured out a way to boost its score on the evaluation by accessing the solutions to what a successful attempt — say, to localize the drone within the 3D map of the office — would look like.
“It’s like an elementary school kid is taking a spelling test and the teacher leaves out an answer rubric on their desk, and then the student squints and tries to look at the rubric,” Sharrock said. “But now, imagine that the student purposely talks to and distracts the teacher so that the teacher forgets that they put the answer key on the desk. That’s what’s happening here.” Sharrock added that newer models tended to cheat in sophisticated ways, a phenomenon highlighted in other recent AI research.
The Andon researchers also discovered that AI systems varied considerably in their willingness to perform the later tasks associated with facial recognition. After an error led to Backlund’s picture not being loaded into a test, Anthropic’s Fable 5 figured it could find another picture by looking for Swedish founders of AI companies, eventually locating and downloading Backlund’s picture.
At the same time, both Anthropic’s Fable 5 and Opus 5 rejected a request to download a picture of an NBC News reporter because of fears of privacy violations and reluctance to use the picture for facial recognition.
The Andon Labs members say they recognize the potential for AI-controlled drones to go wrong, and they insist that Drone-Bench is not meant to encourage people to develop such systems or AI companies to optimize their models to accomplish this behavior.
Instead, they see the project as a critical way to shine light on how anyone, or even AI systems themselves, could soon design and implement facial-recognition tracking technology with cheap consumer drones and off-the-shelf AI models.
“We are not innovating on super cutting-edge autonomous drone tech,” Backlund said. “There are probably thousands of people at many different defense companies doing that without telling the public.”
“We’re doing a very small version of that to tell and show the public where things are heading, and then we should let democracy run its course: Where do we want AI to go, and where should it not go?”