Install
The New Stack is a media platform for the people who build and manage software the world relies on. We provide context and explanation of at-scale technologies to advance knowledge and create conversations through our coverage of modern architectures, components of the software development life cycle, and operations to
- 152articles · 30d
- 6+ hour agolatest article
- Aug 14, 2026earliest in window
- 96%with images
- 88avg words
- Science & Technology 140
- Software Dev. 101
- Computers & Electronics 95
- News 38
- Software 21
- Science & Nature 13
- Economy, Business & Finance 9
- Finance & Business 9
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
“Valuable warning shots”: How Anthropic now views Claude’s cyber incidents
3+ day, 1+ hour ago (857+ words) A new review finds “biased reasoning” and “recklessness” across four Claude cyber incidents, including one its initial search missed....
Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet
3+ day, 7+ hour ago (497+ words) Claude Fable 5.1 more than doubled its predecessor’s benchmark score. On a modest budget, it passed one of five tasks -- but its failures cost less....
Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.
4+ day, 1+ hour ago (492+ words) Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents....
"It could kill us all": what Anthropic's own researchers really think about superintelligence
4+ day, 1+ hour ago (547+ words) ...
Claude Fable 5.1 vs. Fable 5: On real work, I couldn't tell them apart.
1+ week, 1+ day ago (549+ words) Anthropic shipped Claude Fable 5.1 on September 1, claiming doubled performance in agentic research. I ran it against Fable 5 on four real jobs and tracked every token. Both scored perfectly, and on the hardest task, the new model billed more than double…...
Anthropic's Claude failures have made agent observability a security priority
1+ week, 4+ day ago (914+ words) Anthropic aimed to steer its ship into safer, more carefully charted waters this week. The company announced it was improving its alignment and security efforts, and the announcement read somewhat like an admission of responsibility and a mandate for tighter…...
Vercel built a feedback loop that treats agent instructions like software
1+ week, 4+ day ago (391+ words) Vercel cut known design failures by 57% by treating agent guidance like software. Yet none of the six generated pages was ready to ship....
Claude Fable 5.1 watermark: It has a blind spot developers can’t ignore
1+ week, 5+ day ago (303+ words) Anthropic's Fable 5.1 watermarks text but skips code tokens that could break accuracy, while a new API restriction targets model distillation at scale....
Runway wants to generate software as you use it. Solaris is its first step.
1+ week, 5+ day ago (371+ words) The first model of a new class of AI systems Runway calls Interface World Models turns the visual into the application itself....
Anthropic's Fable 5.1 is a bit cheaper, a bit smarter, and refuses a lot less
1+ week, 5+ day ago (594+ words) Anthropic keeps Fable's $10/$50 pricing, cuts cache reads by 75%, and tunes the safeguards that made Fable 5 punt to Opus....