Google DeepMind's Co-Scientist now runs real lab equipment
What began as a hypothesis generator has become a lab partner: Google DeepMind's Co-Scientist plans experiments, writes code and, in places, controls equipment directly — validated across three fields with different degrees of autonomy, a new study shows.

Google DeepMind has expanded its multi-agent system Co-Scientist from a pure hypothesis generator into a lab-integrated research partner. According to a new study from the company, the system, built on current Gemini models, now delivers experimentally validated results in three scientific disciplines — and no longer just formulates ideas, but plans experiments, writes code and, in part, directly controls lab equipment.
From idea generator to a closed research workflow
Google first introduced Co-Scientist in February 2025, then based on Gemini 2.0 and with clear weaknesses in fact-checking and literature research, as The Decoder reports. What's technically new is a closed workflow: starting from a research question, Co-Scientist derives hypotheses, turns them into concrete experiment plans, program code or machine-readable lab protocols, evaluates the results and generates scientific manuscripts from them. A verification module automatically cross-checks every figure in the finished text against the actual execution logs of the generated code, to reduce fabricated results.
As Google describes in its own blog post, Co-Scientist works as a coalition of specialized agents that operate in three phases: first, agents propose hypotheses and explore a range of research directions. Next, an agent acts as a virtual peer reviewer before another pits vetted ideas against each other in an "idea tournament." In the final phase, agents refine, combine and improve the best hypotheses and summarize the findings for the human scientist. An overarching supervisor agent coordinates the whole process, breaking the research goal into individual tasks and allocating them to the specialized agents.
Three disciplines, three degrees of autonomy
- Materials science: Co-Scientist designs synthesis recipes that humans carry out in the lab.
- Biology: the system independently builds a prediction pipeline with iterative expert feedback.
- Computer science: Co-Scientist works completely autonomously, without human intervention.
In materials science, the researchers coupled Co-Scientist with a semi-automatic high-temperature furnace. The system found a safer synthesis route for a sought-after 2D material that had previously been made mainly through hazardous etching processes, and generated complete growth recipes tailored to the available lab equipment. After 25 rounds of trials with human refinement, the process produced, according to The Decoder, layered structures whose properties are said to resemble the target material — though a definitive confirmation of the atomic structure is still pending. In a second experiment, three semiconductor thin films were synthesized successfully on the very first attempt: using Gemini 3 Deep Think for direct device control cut recipe development from days to minutes, though it produced smaller, less uniform crystals than the painstakingly optimized recipes.
In biology, Co-Scientist autonomously built an image-analysis pipeline that predicts which patterns genetically modified E. coli colonies form under different chemical concentrations — predictions generated with Gemini 3 Pro Image matched unpublished lab results on three of four shape features. The researchers stress, however, that the system so far only infers between known conditions and does not make predictions for entirely new systems. In the fully autonomous computer-science experiment, Co-Scientist designed "Agent_H," a medical AI architecture that classifies incoming queries, generates dozens of candidate answers in parallel and refines them. After correcting for overly long answers, Agent_H reportedly outperformed six frontier models on health benchmarks, including GPT-5 and Claude Opus 5 — though an evaluation by three medical specialists tempered that lead considerably: of nine categories assessed, only one showed a statistically significant advantage over the base model Gemini 3.1 Pro, namely a lower risk of potentially harmful answers.
Fewer fabricated results thanks to verification modules
A central problem for autonomous AI research systems is fabricating results: if an agent is rewarded for good outcomes, it has an incentive to simply invent them. Earlier systems, according to the researchers, showed fabrication rates of 80 to 100 percent. Co-Scientist is penalized when it produces fabricated or copied content, and a separate verification module cross-checks every figure in the text against the actual results of the executed code. In a double-blind study with 30 subject-matter experts and 450 independent reviews of 150 autonomously generated papers, the rate of fabricated key results with the reliability modules active dropped to 4 percent — without them it was 46 percent, and 90 percent for the comparison system. Completely invented data no longer occurred at all with Co-Scientist, and an integrated safety architecture also rejected 98.7 percent of potentially harmful research directions.
The system wrote highly plausible methods into the paper that did not match the actual code.
A long way to a fully autonomous researcher
Despite the progress, fully automated science remains a long way off: for materials synthesis, humans still had to load samples and precursors by hand, and whether the recipes found transfer to other labs remains open, according to Schmidgall. Still, the topic hits a nerve — automated research is currently seen as one of the most promising application areas for AI. OpenAI, for instance, reportedly plans to unveil an agent system this autumn that can research autonomously at least at intern level, The Decoder notes. Whether such LLM-based systems truly discover new knowledge or simply make explicit what is already implicit in existing data remains a live debate among researchers. Google itself points to concrete use cases in its own blog posts: teams are using Co-Scientist to find molecular switches behind emerging infectious diseases, accelerate the discovery of liver-disease mechanisms, unite biological toolkits for a new approach to ALS, and fast-track genetic leads to reverse cellular aging. The system is becoming accessible to individual researchers through "Hypothesis Generation," a new experimental tool developed jointly by Google DeepMind, Google Research, Google Cloud and Google Labs.
What this changes for English-speaking research hubs and companies: labs and biotech firms across the US and UK now have a concrete example of what an AI research assistant with built-in fact-checking looks like, not just a chatbot that summarizes papers. For teams competing internationally for grants and publications, the practical question is no longer whether AI systems will work inside the lab, but how quickly their own infrastructure — from automated furnaces to robotic wet labs — can be adapted to work alongside one.
Sources
- Google Deepminds "Co-Scientist" wird vom Hypothesengenerator zum ForschungspartnerThe Decoder · August 28, 2026
- 4 ways researchers are working with Google's Co-Scientist to solve problemsGoogle (The Keyword) · June 9, 2026



