Sõpruse pst, 10615, Tallinn, Estonia +372-55650441

Announcements

No announcements yet

New blog posts

AI Can Run the Experiments — But It Can’t Yet Ask the Right Questions
AI Can Run the Experiments — But It Can’t Yet Ask the Right Questions

17 August, 2026 by Mehrdad Fathi

The promise of a fully autonomous “AI...

Completion of Workshop on Water Recling Simulation and Modelling: Unlocking the Future of Water Management
Completion of Workshop on Water Recling Simulation and Modelling: Unlocking the Future of Water Management

19 March, 2024 by Charlotte Lee

We are thrilled to announce the successful...

IJITIS Journal Meeting and SWOT Analysis at TULTECH
IJITIS Journal Meeting and SWOT Analysis at TULTECH

15 January, 2024 by Charlotte Lee

Greetings, TULTECH community! In our...

View all blog entries →

Journals

 

 

Weather

Clear

18°C

Clear in Tallinn

Calendar of Events

Closest Events
All events on this day

AI Can Run the Experiments — But It Can’t Yet Ask the Right Questions

Posted on 17 August, 2026 by Mehrdad Fathi

AI Can Run the Experiments — But It Can’t Yet Ask the Right Questions

The promise of a fully autonomous “AI scientist” — a system that generates hypotheses, runs experiments, and writes up publishable findings without human intervention — has been one of the most ambitious claims in the technology sector. A new study, reported by Nature this August, delivers a sobering and instructive verdict: even in the field AI knows best, its own, artificial intelligence is not yet ready to research itself.

A Higher Bar Than Peer Review

Researchers at Princeton University, led by computer scientist Sayash Kapoor, designed a rigorous new test they call shadow evaluation. Rather than submitting AI-generated papers to conference peer review — a process Kapoor considers an unreliable quality signal — the team selected two papers submitted to this year’s NeurIPS conference and asked an AI system to independently pursue the same research questions. The original human authors then scrutinized the AI’s output.

The logic is elegant: no reviewer is more qualified, or more motivated, to evaluate a piece of research than the experts who spent months on the same problem themselves.

The system under evaluation was built by harnessing a frontier large language model within an agentic framework, equipped with sub-agents, internet access, software libraries, compute for experiments, and a simulated peer-review tool. Each research task came with a six-day deadline and a $3,000 compute budget.

Where the AI Impressed

The results were not uniformly negative — and the strengths matter for any organization deploying agentic AI today. The system:

  • Ran hundreds of experiments over several days without collapsing into unrecoverable error loops
  • Produced solid literature reviews and generated minor genuine findings
  • Caught its own hallucinations and, contrary to the authors’ expectations, did not attempt to “reward hack” or cut corners

This is a meaningful engineering milestone. Sustained, multi-day autonomous operation with self-correction was not a given even a year ago.

Where It Failed — and Why That Matters

Despite the operational competence, the original authors scored the AI’s research output at 2/6 and 1/6. The failure pattern is worth studying closely, because it mirrors risks that enterprises face when delegating complex work to autonomous systems:

  1. Premature commitment. The system selected a hypothesis early and did not backtrack sufficiently when the approach underperformed.
  2. Insufficient self-criticism. Its internal review loop was too lenient, allowing weak claims to survive — gradually whittled down until little of substance remained.
  3. Poor context awareness. It underused its available time and compute, deviated from instructions, and delivered poorly written, poorly formatted output.

In short: the system executed well but judged poorly. Kapoor’s conclusion is that current AI lacks research creativity — and, critically, the study did not even attempt to evaluate research taste, the ability to decide which questions are worth pursuing in the first place.

Notably, the field is not unanimous. Cong Lu, co-creator of the pioneering “AI Scientist” project, argues that many of these failures were artifacts of an insufficiently constrained harness and could be “trivially fixable” with tighter guardrails — a reminder that system design, not just model capability, shapes outcomes.

The Takeaway for Technology Leaders

For organizations investing in agentic AI, this study offers a practical framework rather than a reason for pessimism. Autonomous systems are already reliable executors: they can sustain long-horizon technical work, self-monitor, and operate within budgets. What they cannot yet do is exercise strategic judgment — knowing when to abandon a failing approach, how to critique their own work honestly, and which problems deserve attention at all.

The operational implication is clear: deploy AI agents where execution is the bottleneck, and keep humans firmly in the loop where judgment is. The Princeton team is already testing improved models and scaffolds, so this boundary will keep moving. But as of today, the most valuable configuration remains a partnership — machine endurance guided by human taste.
 

Source:

Hutson, M. “AI isn’t ready to research itself.” Nature News, 13 August 2026.

https://www.nature.com/articles/d41586-026-02494-5

References:

1.Kirgis, P. et al. Preprint at arXiv https://doi.org/10.48550/arXiv.2607.27191 (2026).

2.Lu, C. et al. Nature 651, 914–919 (2026).


Today In History

Here are some interesting facts ih history happened on 17 August.

  1. Robert Fulton's steamboat Clermont begins 1st trip up Hudson River
  2. US takes LA
  3. 1st bank in Hawaii opens
  4. Federal batteries & ships bombarded Ft Sumter in Charleston
  5. Asaph Hall discovers Mars' moon Phobos
  6. Gold discovered at Bonanza Creek in Klondike region of the Yukon
  7. Bank of Italy opens it's new HQ at Clay & Montgomery
  8. Mob lynches Jewish businessman Leo Frank in Cobb County Ga. after death sentence for murder of 13-year-old girl commuted to life
  9. The Wizard of Oz opens at Loew's Capitol Theater in NY
  10. FDR & Canadian PM William M King agree to joint defense commission
  11. US bombers staged 1st independent raid on Europe attack Rouen France
  12. Allied forces gained completed control of Sicily
  13. Yanks Johnny Lindell ties record with 4 doubles in a game
  14. Indonesia declares independence from the Netherlands (National Day)
  15. Alger Hiss denied ever being a Communist agent
  16. Indonesia gains it's independence
  17. Francis Gary Powers U-2 spy trial opens in Moscow
  18. Gabon gains independence from France (National Day)
  19. Alliance for Progress established
  20. E German border guards shot & mortally wounded Peter Fechter 18 who attempted to cross the Berlin Wall into the western sector
  21. Hurricane Camille claimed more than 250 lives
  22. USSR launches Venera 7 to Venus
  23. 1st manned balloon crossing of the Atlantic Ocean (Eagle II)
  24. Nazi Rudolph Hess dies at 93 after 46 years in Spandau Prison