AI Scientists Are Here—but What Can They Actually Discover?

AI scientists can generate hypotheses, run code and help direct experiments, but discovery requires more than a polished paper. Here is what current systems have actually achieved—and how to judge the evidence.

Candidate hypotheses move through automated laboratory testing to one verified result

Reviewed September 29, 2026. An AI scientist can now search papers, generate hypotheses, write code, plan experiments, analyze results and draft a manuscript. In some tightly defined settings, it can even send instructions to laboratory equipment. That sounds like science without scientists. It is not.

The useful question is narrower: which part of discovery did the system actually complete, and what evidence survived outside the system? A fluent literature summary is not a new hypothesis. A new hypothesis is not an experimental result. A successful experiment is not automatically a durable discovery.

Recent systems have crossed important thresholds. Google’s Co-Scientist helped produce biomedical hypotheses that researchers tested in vitro. FutureHouse’s Robin connected literature search, hypothesis generation and data analysis in a lab-in-the-loop workflow. Sakana AI’s AI Scientist generated machine-learning papers end to end, and one later paper passed workshop review. Self-driving laboratories have also run physical experiments in chemistry and materials science. The evidence is real—but so are the boundaries.

What is an AI scientist?

“AI scientist” is an informal label for a system that coordinates several stages of research rather than performing one isolated task. Most current examples combine a foundation model with search, scientific databases, code execution, specialist models and, sometimes, robotic or cloud-lab interfaces.

The agent architecture matters because science is not one prompt. A system may need to retrieve relevant work, compare competing explanations, rank hypotheses, design a test, operate software or instruments, interpret noisy data and decide what to try next. Multi-agent systems divide those jobs among specialized components that generate, criticize and refine one another’s work.

That makes them more capable than a general chatbot, but the label hides large differences in autonomy. One system may summarize papers and suggest a study. Another may write and run simulation code. A self-driving lab may choose the next experiment inside a human-defined search space. Calling all three “AI scientists” does not mean they have the same evidence, freedom or reliability.

The four levels people often confuse

Level What the AI does What it does not yet prove
Summarization Retrieves and organizes existing findings That the synthesis is complete, correct or novel
Hypothesis generation Proposes a testable explanation, target or candidate That the hypothesis is true
Experimentation Writes code, specifies protocols or controls instruments That the design is valid or the result generalizes
Discovery Contributes to a new, empirically supported finding That the finding will replicate, translate or matter in practice

The distinction is more than semantics. Scientific knowledge accumulates through evidence, criticism, confirmation and correction. A result becomes more trustworthy when methods are transparent and independent researchers can test whether it holds under the same or related conditions.

What current systems have actually discovered

Co-Scientist: hypotheses with laboratory follow-through

Google’s Co-Scientist uses specialized agents to generate, critique, rank and evolve hypotheses from a researcher-defined objective. In work published in Nature, the system helped identify drug-repurposing candidates and combination therapies for acute myeloid leukemia that were then validated in cell experiments.[1]

This counts for more than automated summarization because the output became a falsifiable proposal and met new experimental data. It still does not show that Co-Scientist independently chose the disease, ran the wet lab or established a treatment. In-vitro evidence is an early research result, not a clinical outcome. The scientists supplied objectives, constraints, laboratory work and judgment about what deserved testing.

Robin: an iterative discovery loop, with humans at the bench

Robin links literature agents to a data-analysis agent. In a dry age-related macular degeneration project, it proposed enhancing retinal pigment epithelium phagocytosis, helped select therapeutic candidates and analyzed follow-up data. The published study reports in-vitro efficacy for ripasudil and KL001, then describes a follow-up RNA-sequencing analysis that suggested a possible mechanism and target.[2]

This is one of the clearest examples of AI contributing across multiple intellectual stages: question decomposition, candidate generation, experiment planning, analysis and hypothesis revision. Yet the paper itself calls the process semi-autonomous. Researchers performed the physical experiments and supplied their results. The drug candidates remain preclinical; identifying activity in cultured cells is not the same as showing safety or benefit in patients.

The AI Scientist: producing research is not the same as establishing truth

Sakana AI’s system targets computational machine-learning research. It can propose ideas, edit code, run experiments, make plots, write a paper and review the result. A paper generated by its later system passed the first round of review at an ICLR workshop. The peer-reviewed Nature report notes that the workshop’s acceptance rate was 70% and describes both template-based and template-free modes.[3]

Passing review is a meaningful publishing milestone, but it is not a certificate of scientific truth. An independent evaluation of the earlier system found that five of twelve proposed experiments failed because of coding errors. It also reported weak novelty checks, sparse and dated citations, structural mistakes and hallucinated numerical results.[4] The contrast is instructive: an agent can produce a recognizable paper faster than it can reliably establish that its claim is new, correctly implemented and supported.

Self-driving laboratories: where AI meets physical evidence

In chemistry, Coscientist combined web and documentation search, code execution and laboratory automation. It planned and performed tasks including reaction optimization, although parts of the setup still required manual plate movement and researchers supervised the experiments.[6]

In materials science, A-Lab combined robotics, computational databases, literature-trained models, X-ray diffraction analysis and active learning. The corrected Nature record reports 353 experiments over 17 days and successful synthesis of 36 out of 57 target inorganic materials.[7] This is strong evidence for autonomous experimental optimization inside a specified domain. It is not evidence of an unrestricted machine choosing its own scientific agenda.

These systems expose the central trade-off. The more structured the environment—standardized inputs, machine-readable instruments, objective measurements and a bounded search space—the more autonomy is possible. Open-ended biology, field science and questions involving ambiguous evidence leave far more work for people.

Where AI scientists are strongest today

  • Searching a vast candidate space. Agents can compare more papers, compounds, parameters or experimental paths than one researcher can inspect manually.
  • Turning a goal into a testable shortlist. Ranking candidates can reduce the number of expensive physical experiments.
  • Running repetitive computational work. Code execution, simulation, parameter sweeps and standardized analyses are easier to audit than free-form prose.
  • Closing a bounded experimental loop. When instruments expose reliable software interfaces, a system can use results to choose the next trial.
  • Documenting intermediate work. Tool logs, code and machine-readable protocols can make a workflow easier to inspect—if they are preserved and disclosed.

The near-term opportunity is therefore not a synthetic genius working alone. It is a research stack that compresses search, implementation and iteration while asking humans to spend more attention on framing, anomaly detection, causal interpretation and consequential decisions.

Why the hardest part is still verification

An AI system can generate plausible hypotheses much faster than a laboratory can test them. That creates a verification bottleneck: more candidate claims competing for limited time, equipment, samples and expert attention.

Implementation is one warning sign. PaperBench asks agents to reproduce the empirical contributions of 20 prominent machine-learning papers from scratch. In the published benchmark, the best tested agent averaged a 21% replication score; machine-learning PhDs reached 41.4% on a smaller three-paper subset after 48 hours.[5] These figures are protocol-specific and should not be treated as a universal intelligence ranking. They do show that rebuilding existing research correctly remains difficult—before asking an agent to create reliable new work.

Other failure modes are less visible:

  • Novelty errors. A system may rediscover a known idea because retrieval missed the right terminology, field or negative result.
  • Confirmation bias inside the agent loop. Agents that generate and judge one another’s work may share the same model blind spots.
  • Result hallucination. A polished manuscript can state numbers that were not produced by the experiment or misread a failed run.
  • Proxy optimization. The system may improve the metric it was given while making the real scientific objective worse.
  • Hidden human scaffolding. Choosing the question, building the code template, restricting the search space and translating a plan into a lab protocol can contain much of the scientific judgment.
  • Weak external validity. A result in a benchmark, simulation, cell line or one automated laboratory may not survive a new setting.

A 2025 review of self-driving laboratories concludes that leading systems can automate nearly the entire scientific method in selected settings, while also emphasizing cost, safety, cybersecurity, accountability and intellectual-property challenges.[8] The word selected matters. Autonomy is always autonomy over a particular objective, toolset and environment.

How to judge an “AI discovery” claim

Before accepting the headline, ask seven questions:

  1. What was genuinely new? Was it a summary, a hypothesis, a candidate, an experimental result or a new causal explanation?
  2. Who set the objective and constraints? A broad prompt is different from a carefully engineered search space and human-written scaffold.
  3. What did the AI execute? Separate generated prose from code that ran, instruments that moved and measurements that were collected.
  4. Where did humans intervene? Look for sample preparation, protocol translation, troubleshooting, candidate selection and interpretation.
  5. Was the key claim checked against raw outputs? The manuscript should trace back to logs, data, code and instrument records.
  6. What level of validation was reached? Simulation, in-vitro testing, animal studies, clinical trials and independent replication answer different questions.
  7. Can another team reproduce it? Open code, data, prompts, protocols, model versions and negative results make that possible.

This test prevents two equal mistakes: dismissing a real advance because a human remained in the loop, or calling a generated proposal a discovery before nature has answered.

Will AI replace scientists?

It will replace and reshape parts of scientific work. Literature triage, routine coding, parameter search, first-pass analysis and standardized experimental cycles are obvious candidates. It may also let small teams explore more directions and make advanced automation available through shared cloud labs.

But current evidence favors the co-scientist model. Humans still decide which problems matter, whether a proxy captures the real question, when an anomaly deserves attention, how much evidence is enough and whether a technically possible experiment should be run. Those are not decorative steps around the science; they determine what the science means.

The best AI scientists today are engines for proposing and testing possibilities inside environments we have made legible to machines. Their strongest “discoveries” are joint products of models, databases, instruments and human judgment. That is already consequential. It is also more precise—and more useful—than saying the autonomous scientist has arrived.

Frequently asked questions

Can an AI scientist make a discovery on its own?

In a bounded computational or automated-lab workflow, an AI system can perform many steps with little intervention. Published examples still depend on humans to define goals, build infrastructure, perform or supervise physical work, and validate the result. “On its own” therefore needs a stage-by-stage accounting.

Is an AI-generated paper evidence of discovery?

No. It is evidence that a system can produce a paper. The scientific claim still needs correct implementation, appropriate experiments, transparent data and external scrutiny.

Does peer review prove that an AI result is true?

No. Peer review is a quality filter, not replication. Reviewers can assess novelty, method and presentation, but they generally do not rerun the entire study.

Which fields are most ready for AI scientists?

Fields with digital data, executable models, standardized measurements and programmable instruments are the easiest to automate. Machine learning, chemistry and materials science provide many early examples. Wet-lab biology can support iterative AI workflows, but physical execution and translational validation remain major constraints.

What would count as stronger evidence?

An independently reproduced result, obtained from disclosed methods and raw outputs, would be stronger than a developer demonstration or one-team study. For medical candidates, animal studies and clinical trials would add further levels of evidence; a cell-culture result alone does not establish a treatment.

Disclosure and sources

This article is based on published papers, preprints and official research documentation available on September 29, 2026. It is not a hands-on test of the systems discussed. “Discovery” is used narrowly for a new claim supported by empirical evidence; the strength and maturity of that evidence are stated separately.

  1. Gottweis et al., “Accelerating scientific discovery with Co-Scientist,” Nature (2026).
  2. Ghareeb et al., “A multi-agent system for automating scientific discovery,” Nature (2026).
  3. Lu et al., “Towards end-to-end automation of AI research,” Nature (2026).
  4. Beel, Kan and Baumgart, “Evaluating Sakana’s AI Scientist,” arXiv (revised 2025).
  5. Starace et al., “PaperBench: Evaluating AI’s Ability to Replicate AI Research,” ICML (2025).
  6. Boiko et al., “Autonomous chemical research with large language models,” Nature (2023).
  7. Szymanski et al., “An autonomous laboratory for the accelerated synthesis of inorganic materials,” Nature (corrected 2026).
  8. Tobias and Wahab, “Autonomous ‘self-driving’ laboratories,” Royal Society Open Science (2025).