Patwardhan says OpenAI models took on day-long research tasks
In OpenAI's own DevDay keynote, published 29 September, research lead Tejal Patwardhan gave OpenAI's measurement of its models on day long research tasks.
- In the keynote's research segment the host said "A few weeks ago we reached that goal." and "We now have an AI research intern." (23:52)
- Patwardhan, who introduced herself as "a research lead for post training", said OpenAI's models "have taken on over a third of day long research tasks with no intervention required" (26:27)
- The figure is OpenAI's own measurement of its own models, given as of July; before that, she said, they "failed the vast majority of times" on such tasks (26:27)
OpenAI published its DevDay 2026 keynote on its own YouTube channel on 29 September 2026. The video's description names Sam Altman, Romain Huet, Tejal Patwardhan and Holly Li as the presenters. Introducing the research segment, the host said "Last year we maded a bold prediction saying we would see the first AI research intern in the text year" (23:34), the automatic captions' rendering of what appears to be next year. He then said "A few weeks ago we reached that goal." (23:48) and "We now have an AI research intern." (23:52). Tejal Patwardhan, who introduced herself as "a research lead for post training" (24:17), gave what is OpenAI's own measurement of its own models: they failed most day-long research tasks before, and "as of July and E-our models have taken on over a third of day long research tasks with no intervention required" (26:27). "E-our" is the captions' rendering of a second month name that is not clear. Holly Li, presenting the dots agents, said of OpenAI's own engineers: "Our engineers have their dot fixing dozens of bugs every day." (15:34). Altman also said "Around 1.2 billion people use ChatGPT." (44:11) and, on the API, "We've still kept more than 99 percent reliability." (31:06).
What it means for you Our view
The one-third figure is OpenAI's own measurement of its own models, said on stage, and the claims we hold give neither the list of tasks nor the method behind it. Patwardhan said it failed the vast majority of times on tasks that took a day or longer, so the direction is plain but the size is OpenAI's to defend. The host's "AI research intern" is a description of OpenAI's own work, not an outside verdict. If you are weighing AI for research work, treat this as a vendor claim until an independent test repeats it, and ask the vendor how the tasks were chosen and what "no intervention required" covered.

