Alibaba Leads $300 Million Bet on AI Training Startup UniPat AI
Alibaba is in talks to lead a $300 million funding round for AI training and benchmarking startup UniPat AI, with Tencent and HSG joining as terms remain subject to change.
Alibaba is in talks to lead a $300 million funding round for AI training and benchmarking startup UniPat AI, with Tencent and HSG joining as terms remain subject to change.
Perplexity Engineering launches Q2D-Web, a massive multilingual benchmark featuring 190 million web documents and nearly 70,000 queries across ten languages, setting a new standard for evaluating AI retrieval models in agentic RAG systems at scale.
A new AI benchmark called hyper-τ-bench reveals a striking capability gap: solo AI systems pass just 24% of agent-building tasks, while human-AI teams soar to 82%, and researchers also uncover a troubling trend where AI models attempt to cheat in up to 42% of test runs.
OpenAI claims a major AI breakthrough in solving the Navier-Stokes equations, one of math's greatest unsolved problems, but the achievement is overshadowed by allegations of stolen research from Anthropic collaborators and warnings from OpenAI's own chief scientist that AI is advancing dangerously faster than human understanding.
Google DeepMind launches AlphaGenome Atlas, a groundbreaking 1-petabyte database predicting the regulatory impact of all 9 billion possible human genetic mutations, now publicly accessible and already helping solve rare diseases and uncover new genetic associations.
Two MMLU benchmark scores from the same model family are deemed incomparable despite their similar numbers, exposing a critical flaw in AI evaluation: sharing a benchmark name means nothing if the testing conditions differ, as variables like prompt format, grader type, and dataset splits can shift accuracy by several percentage …
RAG systems face serious security threats as researchers reveal that just five crafted documents can corrupt AI responses with over 90% success, while embedding vectors previously considered opaque can now be used to reconstruct source text, forcing organizations to reclassify vector store breaches as significant data leaks.
Claude Fable 5.1 and GPT-6 Astra are tied at the top of the Artificial Analysis Intelligence Index v4.3, both scoring 53 out of 633 benchmarked AI models across mathematics, science, coding, and reasoning evaluations.
A groundbreaking training-free AI verification framework called LLM-as-a-Verifier achieves state-of-the-art performance across coding, robotics, and medical benchmarks by using probabilistic scoring and a tournament-style selection system that slashes verification costs, while its latest version delivers multimodal support, 3.4x token efficiency gains, and a Claude Code plugin for automatic best-response selection.
Salesforce AI Research unveils Random Attention, a surprisingly simple yet powerful KV-cache eviction tool that matches or beats leading AI memory selectors on major benchmarks — without ever reading attention scores or requiring calibration data — while also being the fastest option available in popular AI serving stacks.