August 12, 2026
Databricks Network Configuration delivery to Tens of Millions of Serverless VMs
SummaryDatabricks' serverless platform launches tens of millions of VMs daily, and...
Read original articleDataAIHub Daily
Archive →39 curated AI news stories from leading AI companies.
August 12, 2026
SummaryDatabricks' serverless platform launches tens of millions of VMs daily, and...
Read original articleAugust 12, 2026
"Every new trainset is a data-generating asset. The intelligence platform that connects...
Read original articleAugust 12, 2026
xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5. On agentic tasks, it completes complex workflows in about 53 steps where Claude Opus 5 needs 103, at a price more than 60 percent lower. The article SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price appeared first on The Decoder.
Read original articleAugust 12, 2026
Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open...
Read original articleAugust 12, 2026
Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Altimeter Capital.
Read original articleAugust 12, 2026
Originally designed to make software predictable for humans, Google Go is now positioning itself as a language tailored for machine The post Code that passes every test can still break the next AI agent that touches it appeared first on The New Stack.
Read original articleAugust 12, 2026
During AI DevCon in London this summer, Lamis Mukta, member of technical staff at Anthropic, hosted a stage presentation session The post Anthropic gave agents the ability to dream. Then developers woke up. appeared first on The New Stack.
Read original articleAugust 12, 2026
One of Canada’s largest railway networks spans roughly 20,000 route miles across...
Read original articleAugust 12, 2026
August 12, 2026
AI infrastructure spans multiple layers, from compute and networking to storage, orchestration, and applications. When performance degrades, identifying the...
Read original articleAugust 12, 2026
Three data sources tell the same story: Google's Gemini is losing AI market share. Pangram reports a drop from 12 to 1.9 percent, while OpenAI holds over 50 percent, and Anthropic grew from 4.3 to 14.9 percent. Similarweb and OpenRouter confirm the trend. The article Google's Gemini is losing market share to ChatGPT and Claude according to new market data appeared first on The Decoder.
Read original articleAugust 12, 2026
From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event.
Read original articleAugust 12, 2026
August 12, 2026
Google’s new Pixel 11 puts Gemini at center of AI phone battle with Apple CNBC
Read original articleAugust 12, 2026
From Magic Capture and Instant Night Sight to a built-in teleprompter, here’s a look at a few camera features on Google’s new Pixel 11 series.
Read original articleAugust 12, 2026
Robert Mahari is Anthropic's first "Head of Claude for Legal," responsible for deploying and expanding the Claude AI model across the legal industry. The article Legal startup founder Robert Mahari joins Anthropic to lead Claude's push into law practices appeared first on The Decoder.
Read original articleAugust 12, 2026
The Data Analyst role has been declared dead more times than we can count. AI will...
Read original articleAugust 12, 2026
Meta released Muse Glimmer on Monday, a 30-billion-parameter open-weight model distilled from Muse Spark and licensed under Apache 2.0. The The post Meta stopped worrying about distillation and just shipped the pipeline appeared first on The New Stack.
Read original articleAugust 12, 2026
Nvidia is working on Nemotron 4, a new open-weight model designed to rival the world’s best freely available models. The article Nvidia's Nemotron 4 aims for one trillion parameters, a scale Chinese labs already surpassed appeared first on The Decoder.
Read original articleAugust 12, 2026
Meta and Nvidia plant 'very firm flag' in open-weight AI race led by Chinese Labs CNBC
Read original articleAugust 12, 2026
Judges around the world have made headlines for illicitly using generative AI in their work. But in Pakistan, a large-scale trial of a specially designed AI tool for judges found the technology–together with appropriate training–boosted the number of cases resolved by 6.3% with no obvious drop in the quality of judgements.With a backlog of 2.26 million cases and fewer than two judges per 100,000 people – compared to 22 in the EU and 8 in Brazil – Pakistan’s judiciary was in sore need of help. So, in consultation with the judiciary, economist Sultan Mehmood, of the New Economic School in Moscow, Russia, and collaborators tested whether AI could ease the burden.They built a custom tool combining OpenAI’s GPT-4 large language model (LLM) with a knowledge base of nearly 130,000 Pakistani judicial opinions and statutes, to help judges with legal research and drafting judgements. They began offering the tool in 2024 to 1,559 trial judges – roughly half the country’s justices. “We do find an increase in cases resolved, and we don’t find any corresponding decrease in decision quality,” Mehmood says. First of its kind“It’s pretty amazing that he’s able to pull this off,” says David Autor, an economics professor at the Massachusetts Institute of Technology. “It’s not easy to do large-scale field experiments in civil service, but especially where the stakes are so high.” The 6.3% productivity boost is not overwhelming, he says, but it’s credible and likely to improve as the tool is more widely used.AI tools for judges are already being rolled out in Brazil and India, and prominent U.S. law professor Eric Posner has compared LLM judgments to human judgements in a single case study, but until now there have been no major independent assessment of ongoing judicial use of AI. The new study focused on Pakistan’s trial courts; Mehmood says judges there were enthusiastic from the start.“They were more techno-optimist than we were,” he says. “The delays are so huge this is something which they thought was worth trying anyway to reduce people’s suffering.”Some judges were also already using AI chatbots, Mehmood says, but commercial offerings performed poorly on Pakistani legal queries, frequently hallucinating case law. So the team built a tool tailored to the Pakistani context, called JudgeGPT.They used retrieval-augmented generation (RAG), which allowed the model to query a database of 128,292 Pakistani judicial opinions and 943 statutes. Responses included footnotes linking to cases and laws.“It turns out that actually the way to fix [hallucinations] isn’t just more intelligent models,” says study co-author Elliott Ash, an associate professor of law, economics, and data science at ETH Zurich in Switzerland. “It’s to attach the models to a tool that can do a search and verify the sources.” However, the researchers do not report hallucination rates.The team also put 1,197 judges through six 90-minute Zoom training sessions where Ash covered how LLMs work, their limitations, the risk of bias and hallucinations and the importance of verifying outputs. Another 180 judges only underwent general training on technology in legal research, while a final group got no training.By the time 487 judges had been through the program, the median district saw a jump of 6.3% resolved cases, and the more trained judges in a district, the bigger the effect. Appeal rates also fell slightly, suggesting faster resolution wasn’t leading to sloppier decisions. The JudgeGPT can be used to surface relevant case law with a simple text query and results provide links to the full text of the related judgements.Sultan Mehmood, Christoph Goessmann, and Elliott AshAddressing limitationsThe team also assessed the quality of judgements. Having legal experts evaluate large numbers of judgements was infeasible, Mehmood says, so the team asked OpenAI’s GPT-5-mini to choose between pairs of judgements from the same judge before and after training. The LLM chose post-training judgements 59% of the time. Two experienced Pakistani lawyers also evaluated the model’s analysis of 90 judgement pairs. They agreed with GPT-5-mini 70.6% of the time, compared to 73% agreement with each other.Training turned out to be vital. On average, JudgeGPT-trained judges logged in 56 times and sent 212 prompts over the study period, compared to 10 logins and 25 prompts after generic training. Those who had no training tended to use the tool for around a month and then drop off entirely, Mehmood says. “Just giving people the technology does not necessarily make them use it persistently,” he says.A 6.3% increase sounds modest, but the researchers calculated that a trained judge was resolving 38.5 more cases a month than the baseline, translating to roughly US $38.50 saved in judicial costs for every dollar spent running the tool. Ash also notes that these figures come from a nine-month period at the start of the trial, and that they’ve since updated both the underlying AI model and the database.For users the tool has been a lifeline. One participating trial judge, who spoke on condition of anonymity, says the number of cases assigned to them hasn’t dropped below 1,000 in more than a decade. The tool saves significant time, in particular searching for case law and summarising lengthy documents. “For research, it’s just one prompt away, whereas before I had to search for the precedents and laws for hours,” the judge says. “If I have to read 10 pages of a precedent, now I ask JudgeGPT to just summarise it for me and give me the crux, and it does that work in seconds.”But efficiency isn’t the only thing you want out of a justice system, says John Zeleznikow, professor of law and technology at La Trobe University in Australia. “What they’ve tried to do is be effective, [to] deal with more cases more quickly, and they’re able to do that,” he says. “What’s not that clear is whether what you call the quality of justice is better.”Zeleznikow saysAI can be useful but only if judges are diligent about evaluating and verifying the output. However, the working paper’s authors found that roughly a fifth of participant’sthe prompts given to JudgeGPT involved what they call “substantial AI delegation” – asking the tool what the best decision is, to produce legal reasoning or write opinions with little input from the judge. On the bright side, training lowered the proportion of inappropriate delegation.But given that judges are already using AI, Ash says better tools and training are crucial. “There are risks for using these AIs for sure, even with all these safeguards, but at some point you have to just put the judges in as strong a position as you can,” he says. “Have technological safeguards, but then try to encourage the judges not to rely on it too much.”
Read original articleAugust 12, 2026
I Asked AI to Write a Novel. It’s Not So Bad. Mother Jones
Read original articleAugust 12, 2026
Last week, an AI newsroom beat mainstream journalists—including WIRED—to a story about OpenAI and hacking. It’s just the beginning.
Read original articleAugust 12, 2026
Microsoft has released MAI Code 1.1 Flash, a code model for GitHub Copilot that's said to be 25 percent more token-efficient at a quarter of the cost of its predecessor. In benchmarks, though, it gets crushed by the cheaper Deepseek V4 Flash. The move fits a pattern: Microsoft talks up open AI, then bakes worse proprietary models into its apps to protect margins. The article Microsoft's new MAI Code 1.1 Flash gets crushed by Deepseek on both price and performance appeared first on The Decoder.
Read original articleAugust 12, 2026
Zuckerberg lays out Meta’s new vision for AI: Personal intelligence for all Los Angeles Times
Read original articleAugust 12, 2026
NVIDIA's open 30B MoE targets the agent execution layer, with Switchyard routing each step to the cheapest capable model. The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.
Read original articleAugust 12, 2026
OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.
Read original articleAugust 12, 2026
arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G\"odel Machine and the Huxley G\"odel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modification search over many candidate agents. For planning tasks with explicit constraints, we propose a far cheaper alternative. We introduce SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer in the same closed-loop, improve-from-experience family as the G\"odel-machine methods, but self-supervised rather than self-referential. Given an agentic harness, SBCO learns a decomposed bank of verifiers and a harness policy via approximate block coordinate ascent, improving the agent's outputs from its own graded feedback---with a fixed meta-agent and no human labels. Across two domains SBCO matches or exceeds a customized self-modifying baseline while using 4-5.5 times less compute budget.
Read original articleAugust 12, 2026
arXiv:2608.10008v1 Announce Type: new Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. We jointly audit hallucination rate (OOD@10) and verbalized-confidence calibration (ECE, Brier, reliability) for four zero-shot LLM recommenders from four independent vendors (Mistral Large, Llama-3.3-70B, GPT-OSS-120B, Claude Sonnet 4.6), not grounded or fine-tuned systems, across three catalogs (MovieLens-25M, Amazon Reviews 2023 Toys, Yelp Open Dataset), stratified by item popularity. Hallucination is catalog-dependent (0--0.2\% on MovieLens, 4.5--8.3\% on Amazon, 2.2--8.4\% on Yelp), but verbalized confidence is materially miscalibrated even when hallucination is zero (ECE up to 0.223 on MovieLens despite 0\% OOD). All four LLMs are systematically \emph{under}-confident across all twelve cells, verbalizing a mean of 67--86 on items they recommend with 92--100\% accuracy. This is the opposite of the over-confidence usually emphasized in LLM-hallucination work. The under-confidence is best read as an \emph{elicitation mismatch}: ``Just Ask'' elicits a generic recommendation-quality rating, not a catalog-membership probability. A conformal abstention threshold over verbalized confidence reduces hallucination by at most 0.7\,pp across $\alpha \in \{.05, .10, .15, .20\}$, at 4--21\,pp of coverage cost: the under-confident channel cannot separate correct items from hallucinations, so the threshold mostly removes correct items. We recommend that audits of LLM recommenders report calibration alongside OOD, and use catalog-anchored elicitation rather than generic confidence prompts.
Read original articleAugust 12, 2026
arXiv:2608.10187v1 Announce Type: new Abstract: Modern recommendation systems on social media platforms such as Meta must model complex social relationships, including friendships, group memberships, and creator interactions, alongside massive and heterogeneous content such as text and video. Traditional recommendation models, however, often omit these signals or treat them independently, lacking the reasoning capability to integrate multi-relational context for fine-grained personalization. We present ConnectionMind, a production-ready recommendation framework that tightly integrates the social network structure with large language models (LLMs) to enable scalable, interpretable, and reasoning-aware personalization in Meta. ConnectionMind constructs a heterogeneous graph connecting users, items, friends, groups, and creator pages, and formulates recommendation as a graph reasoning problem: discovering personalized paths from users to candidate items. An LLM-based policy is employed to reason over these graph structures and guide recommendation decisions. To train the system at scale, ConnectionMind adopts a two-stage learning strategy. We first perform supervised fine-tuning (SFT) on large-scale user-item interaction trajectories to initialize the reasoning policy, followed by end-to-end reinforcement learning (RL) to refine the model's ability to reason over social graphs for personalized recommendation. Extensive experiments on multiple real-world datasets demonstrate the effectiveness of ConnectionMind compared to representative baselines. More importantly, ConnectionMind has been deployed in Meta's large-scale recommendation pipeline and has been evaluated through online A/B tests, achieving a 0.43% improvement in video watch time. These results demonstrate measurable real-world impact in a production recommendation system.
Read original articleAugust 12, 2026
arXiv:2608.10385v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.
Read original articleAugust 12, 2026
Artificial Intelligence in Gallbladder Imaging: A Rapid Evidence Review and Exploratory Meta-Analysis of Diagnostic Performance and Reader Assistance Cureus
Read original articleAugust 12, 2026
Nvidia’s Show of Financial Force Soothes Jittery Credit Markets Bloomberg.com
Read original articleAugust 12, 2026
Day-0 vLLM support for Qwen3.8-2.4T-A95B: a 2.4-trillion-parameter hybrid MoE model served out of the box, with FP8/BF16 checkpoints plus NVFP4 and MXFP4 quantized weights, and co-developed kernels on NVIDIA and AMD hardware.
Read original articleAugust 11, 2026
Change data capture is one of the most common things data engineers build on Spark...
Read original articleAugust 11, 2026
River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised $1.1 billion in early-stage funding. The company stated in today’s announcement that it received the capital over two rounds, a seed and a Series A. General Catalyst and AMP PBC were the lead investors. They were joined by Nvidia Corp., […] The post Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD appeared first on SiliconANGLE.
Read original articleAugust 11, 2026
Google LLC’s Gemini artificial intelligence app has passed 1 billion monthly active users, making it the 14th product in the company’s history to reach that mark. The company announced the milestone today in a blog post from Josh Woodward, vice president of Google Labs, Gemini and AI Studio. Chief Executive Sundar Pichai said in a […] The post Google’s Gemini AI app passes 1 billion monthly active users appeared first on SiliconANGLE.
Read original articleAugust 11, 2026
Manus is becoming an independent company again after Chinese regulators ordered Meta to unwind its roughly $2 billion acquisition of The post Your AI agent remembers everything. Here’s what happens when its owner changes. appeared first on The New Stack.
Read original articleAugust 11, 2026
Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock | Artificial Intelligence Amazon Web Services (AWS)
Read original article