Sõpruse pst, 10615, Tallinn, Estonia +372-55650441

Announcements

No announcements yet

New blog posts

Can Pancreatic Cancer Be Stopped in Its Tracks? A Radical Tactic Raises Hopes
Can Pancreatic Cancer Be Stopped in Its Tracks? A Radical Tactic Raises Hopes

4 October, 2026 by Mehrdad Fathi

With prospects for new treatments at an...

When AI Paints the Picture: The Microscopy Image Controversy Shaking Science
When AI Paints the Picture: The Microscopy Image Controversy Shaking Science

2 October, 2026 by Mehrdad Fathi

A prize-winning video of lung tissue has...

How to Stay Smart in the Age of AI: The Science of Critical Thinking
How to Stay Smart in the Age of AI: The Science of Critical Thinking

28 September, 2026 by Mehrdad Fathi

There is growing concern that AI can blunt...

View all blog entries →

Journals

 

 

Weather

Clouds

14°C

Clouds in Tallinn

Calendar of Events

Closest Events
All events on this day

AI Can Run the Experiments — But It Can’t Yet Ask the Right Questions

Posted on 17 August, 2026 by Mehrdad Fathi

AI Can Run the Experiments — But It Can’t Yet Ask the Right Questions

The promise of a fully autonomous “AI scientist” — a system that generates hypotheses, runs experiments, and writes up publishable findings without human intervention — has been one of the most ambitious claims in the technology sector. A new study, reported by Nature this August, delivers a sobering and instructive verdict: even in the field AI knows best, its own, artificial intelligence is not yet ready to research itself.

A Higher Bar Than Peer Review

Researchers at Princeton University, led by computer scientist Sayash Kapoor, designed a rigorous new test they call shadow evaluation. Rather than submitting AI-generated papers to conference peer review — a process Kapoor considers an unreliable quality signal — the team selected two papers submitted to this year’s NeurIPS conference and asked an AI system to independently pursue the same research questions. The original human authors then scrutinized the AI’s output.

The logic is elegant: no reviewer is more qualified, or more motivated, to evaluate a piece of research than the experts who spent months on the same problem themselves.

The system under evaluation was built by harnessing a frontier large language model within an agentic framework, equipped with sub-agents, internet access, software libraries, compute for experiments, and a simulated peer-review tool. Each research task came with a six-day deadline and a $3,000 compute budget.

Where the AI Impressed

The results were not uniformly negative — and the strengths matter for any organization deploying agentic AI today. The system:

  • Ran hundreds of experiments over several days without collapsing into unrecoverable error loops
  • Produced solid literature reviews and generated minor genuine findings
  • Caught its own hallucinations and, contrary to the authors’ expectations, did not attempt to “reward hack” or cut corners

This is a meaningful engineering milestone. Sustained, multi-day autonomous operation with self-correction was not a given even a year ago.

Where It Failed — and Why That Matters

Despite the operational competence, the original authors scored the AI’s research output at 2/6 and 1/6. The failure pattern is worth studying closely, because it mirrors risks that enterprises face when delegating complex work to autonomous systems:

  1. Premature commitment. The system selected a hypothesis early and did not backtrack sufficiently when the approach underperformed.
  2. Insufficient self-criticism. Its internal review loop was too lenient, allowing weak claims to survive — gradually whittled down until little of substance remained.
  3. Poor context awareness. It underused its available time and compute, deviated from instructions, and delivered poorly written, poorly formatted output.

In short: the system executed well but judged poorly. Kapoor’s conclusion is that current AI lacks research creativity — and, critically, the study did not even attempt to evaluate research taste, the ability to decide which questions are worth pursuing in the first place.

Notably, the field is not unanimous. Cong Lu, co-creator of the pioneering “AI Scientist” project, argues that many of these failures were artifacts of an insufficiently constrained harness and could be “trivially fixable” with tighter guardrails — a reminder that system design, not just model capability, shapes outcomes.

The Takeaway for Technology Leaders

For organizations investing in agentic AI, this study offers a practical framework rather than a reason for pessimism. Autonomous systems are already reliable executors: they can sustain long-horizon technical work, self-monitor, and operate within budgets. What they cannot yet do is exercise strategic judgment — knowing when to abandon a failing approach, how to critique their own work honestly, and which problems deserve attention at all.

The operational implication is clear: deploy AI agents where execution is the bottleneck, and keep humans firmly in the loop where judgment is. The Princeton team is already testing improved models and scaffolds, so this boundary will keep moving. But as of today, the most valuable configuration remains a partnership — machine endurance guided by human taste.
 

Source:

Hutson, M. “AI isn’t ready to research itself.” Nature News, 13 August 2026.

https://www.nature.com/articles/d41586-026-02494-5

References:

1.Kirgis, P. et al. Preprint at arXiv https://doi.org/10.48550/arXiv.2607.27191 (2026).

2.Lu, C. et al. Nature 651, 914–919 (2026).


Today In History

Here are some interesting facts ih history happened on 4 October.

  1. 1st code of law for Plymouth Colony
  2. Peter Stuyvesant establishes Americas 1st volunteer firemen
  3. Mexico becomes a republic
  4. Market Street's "Path of Gold" lit for 1st time
  5. the dahlia is officially designated as SF city flower
  6. The comic strip Dick Tracy debuts
  7. Wrestling returns to Madison Sq Garden after 12 year lay off
  8. Indians beat Red Sox in 1st AL playoff game
  9. Brooklyn Dodgers only World Series victory beating Yankees
  10. USSR launches Sputnik I the 1st artificial earth satellite
  11. Leave It to Beaver debuts on CBS
  12. Transatlantic coml jet passenger service began
  13. USSR Luna 3 sent back 1st photos of Moon's far side
  14. Courier 1B Launched; 1st active repeater satellite in orbit
  15. Whitey Ford's world series 33 2/3 scoreless inning streak ends
  16. Lesotho (Basutoland) gains independence from Britain (National Day)
  17. UN starts issuing postage stamps at Geneva headquarters
  18. Janis Joplin dies at 27
  19. Agriculture Secretary Earl Butz resigns after racial 'joke'
  20. pier 39 opens in SF
  21. Yanks clinch AL East
  22. 21st Space Shuttle Mission - Atlantis 1 is launched
  23. James Jefferson of Winnipeg scores 2 TDs on interception returns without making an interception. (He scored on laterals.)