DeepSeek's V4 Flash tops AI leaderboards but completed just 53.8% of real-world agent tasks in a new test — as the company ...
That last finding is the one that qualitative review would never have surfaced. The model's expressed confidence didn't ...
Credit: VentureBeat made with OpenAI ChatGPT-Images-2.0 Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today ...
For enterprise developers, Harness may ultimately be the more consequential part of Thursday’s announcement. Models can ...
The headline numbers are striking: Writer says its agent product now operates at an average 52% lower cost, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6 ...
Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models ...
Four of five enterprises that secured AI agent identities never built isolation to contain a compromised agent, VentureBeat's July Pulse survey found.
Grok 4.6 does not establish an uncontested performance lead. Its launch instead presents a different proposition: frontier-level intelligence, large improvements over the previous generation, stronger ...
Google is rolling out Gemini 3.7 Flash, a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API ...
SLA-backed regional inference, five-year European Compute Unit contracts, hosting China's GLM-5.2, and 1 gigawatt of compute ...
A VentureBeat survey of 101 enterprises found 68% traced a confident, wrong AI agent answer to bad context. Governed semantic ...
K2 Global backs frontier technology companies including Neuralink, xAI, SpaceX, Shield AI, Tenstorrent, Synchron, Lambda, ...