
01AI mathematics
OpenAI released an AI-generated proof that a smoothly forced three-dimensional fluid can develop an infinite-speed singularity in finite time. The 166-page argument comes with a public Lean formalization. That makes the logic machine-checkable, but mathematicians still need to confirm that the formal statement and interpretation fully match the famous problem.
Why it mattersA machine-generated result on a 90-year-old Millennium problem is an unusually strong test of whether frontier systems can produce new, checkable mathematics rather than only summarize it.
Keep in mind: The mathematical community has not yet completed independent expert review. Formal checking validates the encoded statements and proof steps, not whether the formalization captures every intended interpretation.
Read the source · OpenAI ↗02AI for biology
Google DeepMind opened a one-petabyte atlas containing predicted effects for nine billion possible single-letter DNA changes. Scientists can browse it or query an API instead of running each prediction separately. The entries are model outputs, not laboratory measurements, although collaborators report experimental confirmation for selected rare-disease variants.
Why it mattersScientists can rank candidate variants and inspect predicted mechanisms across coding and non-coding DNA without first running billions of separate model queries.
Keep in mind: Most entries are predictions, not laboratory measurements. Benchmark and collaborator-result framing comes from Google DeepMind.
Read the source · Google DeepMind ↗03AI in the lab
An MIT researcher connected GPT-5.6 Sol through Codex to the software controlling a six-qubit chip. The agent chose parameters, ran measurements, and completed routine calibration sequences when signals were clear. Weak or noisy signals still needed an experienced researcher, so this is useful lab assistance rather than unsupervised discovery.
Why it mattersA researcher can leave well-defined measurements running for hours while reserving human attention for ambiguous signals and experiment design.
Keep in mind: The evidence is a vendor case study rather than an independent evaluation. The agent handled routine calibration, not unsupervised scientific discovery.
Read the source · OpenAI and MIT EQuS ↗04Image making
OpenAI launched a new image model with more reliable targeted edits, steadier multi-step refinement, and up to 50 percent lower generation latency than its previous version. It is rolling out across ChatGPT, Work, Codex, and the API. The quality and speed comparisons come from OpenAI and early partners, not independent testing.
Why it mattersPeople can refine a picture over several turns without rebuilding it each time, while applications gain a faster default model and a slower precision option.
Worth doing: Test a familiar multi-step edit before changing a production image workflow.
Keep in mind: Quality and latency comparisons are vendor-reported; independent testing was not available before the cutoff.
Read the source · OpenAI ↗05AI for chemistry
TSBench asks AI agents to propose three-dimensional transition states, then checks the structures with quantum chemistry instead of grading their prose. Across 546 evaluations, diagnosis-led retries raised aggregate success from 50.4 to 66.8 percent. The remaining failures show that plausible-looking chemistry can still follow the wrong reaction path.
Why it mattersScientific agents can sound chemically plausible while following the wrong reaction path, and this test exposes that failure before an autonomous workflow trusts it.
Keep in mind: This is a preprint benchmark, not proof that the tested agents can plan useful laboratory syntheses.
Read the source · Xiaohu Xu and Tong Zhu ↗06AI for health
Researchers trained SleepFM-2 on more than two million hours from 235,865 overnight recordings, then tested transfer across diseases, expert-scored sleep events, wearables, and subjective reports. The scale is unusual and the held-out tests are broad. It remains a preprint and does not establish clinical benefit or diagnostic approval.
Why it mattersOne learned view of overnight brain, heart, muscle and breathing signals may support many downstream health studies instead of a separate model for each sensor and task.
Keep in mind: The paper is a preprint and does not establish clinical benefit or diagnostic approval.
Read the source · SleepFM-2 authors ↗07Reality check
PhysWeep measures the physical values that actually appear in generated videos, rather than asking whether motion merely looks believable. Two of three open generators repeatedly settled on seed-dependent but incorrect values. The audit covers only three models and six parameter sweeps, but it exposes a clean gap between visual plausibility and physical control.
Why it mattersA clip can look physically natural while ignoring the number or condition a user asked for, which weakens claims that today’s video generators understand physics.
Keep in mind: The audit covers three open generators and six parameter sweeps, not every video model or physical behavior.
Read the source · Rasul Khanbayov and Hasan Kurban ↗