Credit: VentureBeat made with OpenAI ChatGPT-Images-2.0 Meta today released Muse Code, a terminal-based AI coding agent now in beta, alongside Muse Spark 1.2, a coding-focused update to its Muse Spark ...
The Trump administration has reportedly considered restrictions on Chinese AI models after Moonshot AI’s 2.8-trillion-parameter Kimi K3 topped a major coding leaderboard and drew growing interest from ...
Moonshot AI’s Kimi K3 topped a frontend coding benchmark, beating Claude Fable 5 while adding pressure on US AI leaders. Chinese AI startup Moonshot AI has released Kimi K3, a massive open-weight AI ...
OpenAI rolled out three separate GPT-5.6 tiers, Sol, Terra and Luna, replacing its previous single-model approach. Independent testing found Claude Fable 5 edges GPT-5.6 Sol on creative and coding ...
Most widely cited AI coding benchmarks, including the original SWE-bench, were built primarily around Python repositories, meaning headline performance results may not accurately predict how coding ...
AI coding benchmark scores that labs, enterprises, and investors use to compare frontier models are inflated by answer retrieval — not genuine reasoning — and the smarter the model, the more inflated ...
Google’s Android Bench results show Gemini 3.5 Flash trailing older models despite its premium positioning. Gemini 3.5 Flash missed the top five, while OpenAI’s GPT 5.5 claimed first place and Gemini ...
You're currently following this author! Want to unfollow? Unsubscribe via the link in your email. Update June 16, 2026: A day after this story was published, SpaceX announced it would buy Cursor for ...
Lovable, the European "vibe coding" startup that lets people build software by describing it in plain language, told TechCrunch on June 9, 2026, that it has surpassed $500 million in annualized ...
A monthly overview of things you need to know as an architect or aspiring architect. Unlock the full InfoQ experience by logging in! Stay updated with your favorite authors and topics, engage with ...
The latest flare-up in the debate over AI-assisted coding did not come from a new model release or a benchmark result. It came from a single line of text buried inside a software update. Earlier this ...
Anthropic launched Claude Opus 4.8 on May 28, pricing it level with the prior 4.7 release. The company says it tops OpenAI's GPT-5.5 and Google's Gemini 3.1 Pro on SWE-Bench Pro and other tests. A ...