DataAIHub Daily

Archive →

August 12, 2026

39 curated AI news stories from leading AI companies.

Databricks

August 12, 2026

Databricks Network Configuration delivery to Tens of Millions of Serverless VMs

SummaryDatabricks' serverless platform launches tens of millions of VMs daily, and...

Read original article

Databricks

August 12, 2026

How Amtrak is building the data backbone for its largest transformation in over 50 years

"Every new trainset is a data-generating asset. The intelligence platform that connects...

Read original article

Anthropic

August 12, 2026

Space XAI's Grok 4.6 matches Open AI's best model and undercuts it on price

xAI's Grok 4.6 scores 61 points on the Artificial Analysis Intelligence Index, tying GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5. On agentic tasks, it completes complex workflows in about 53 steps where Claude Opus 5 needs 103, at a price more than 60 percent lower. The article SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price appeared first on The Decoder.

Read original article

Nvidia

August 12, 2026

Serve Qwen3.8-2.4 T-A95 B, a 2.4 T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

Alibaba released the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model, bringing near-frontier capabilities to the open...

Read original article

OpenAI

August 12, 2026

Open AI-backed Thrive Holdings raises $2 B to bring AI to the enterprise

Thrive Holdings has raised $2 billion in new funding at a $12 billion valuation from investors like SoftBank, D1 Capital Partners, and Altimeter Capital.

Read original article

Google

August 12, 2026

Code that passes every test can still break the next AI agent that touches it

Originally designed to make software predictable for humans, Google Go is now positioning itself as a language tailored for machine The post Code that passes every test can still break the next AI agent that touches it appeared first on The New Stack.

Read original article

Anthropic

August 12, 2026

Anthropic gave agents the ability to dream. Then developers woke up.

During AI DevCon in London this summer, Lamis Mukta, member of technical staff at Anthropic, hosted a stage presentation session The post Anthropic gave agents the ability to dream. Then developers woke up. appeared first on The New Stack.

Read original article

Databricks

August 12, 2026

How a major freight railroad scaled pipeline creation with Genie Code

One of Canada’s largest railway networks spans roughly 20,000 route miles across...

Read original article

Nvidia

August 12, 2026

How to Choose Full-Stack Observability for NVIDIA AI Factories

AI infrastructure spans multiple layers, from compute and networking to storage, orchestration, and applications. When performance degrades, identifying the...

Read original article

Anthropic

August 12, 2026

Google's Gemini is losing market share to Chat GPT and Claude according to new market data

Three data sources tell the same story: Google's Gemini is losing AI market share. Pangram reports a drop from 12 to 1.9 percent, while OpenAI holds over 50 percent, and Anthropic grew from 4.3 to 14.9 percent. Similarweb and OpenRouter confirm the trend. The article Google's Gemini is losing market share to ChatGPT and Claude according to new market data appeared first on The Decoder.

Read original article

Google

August 12, 2026

Everything announced at Made by Google ’26: Pixel 11, Pixel Watch 5, Pixel Tag, and tons of Gemini features

From the Pixel 11 series and a brand new competitor to Apple’s AirTag, here are all the announcements from the Made by Google 2026 event.

Read original article

Google

August 12, 2026

Google’s new Pixel 11 puts Gemini at center of AI phone battle with Apple - CNBC

Google’s new Pixel 11 puts Gemini at center of AI phone battle with Apple CNBC

Read original article

Google

August 12, 2026

4 New Camera Tricks on Google’s Latest Pixel 11 Smartphones

From Magic Capture and Instant Night Sight to a built-in teleprompter, here’s a look at a few camera features on Google’s new Pixel 11 series.

Read original article

Anthropic

August 12, 2026

Legal startup founder Robert Mahari joins Anthropic to lead Claude's push into law practices

Robert Mahari is Anthropic's first "Head of Claude for Legal," responsible for deploying and expanding the Claude AI model across the legal industry. The article Legal startup founder Robert Mahari joins Anthropic to lead Claude's push into law practices appeared first on The Decoder.

Read original article

Databricks

August 12, 2026

The Future of Data Analytics: Why AI is rewriting the Analyst’s Job Description

The Data Analyst role has been declared dead more times than we can count. AI will...

Read original article

Meta

August 12, 2026

Meta stopped worrying about distillation and just shipped the pipeline

Meta released Muse Glimmer on Monday, a 30-billion-parameter open-weight model distilled from Muse Spark and licensed under Apache 2.0. The The post Meta stopped worrying about distillation and just shipped the pipeline appeared first on The New Stack.

Read original article

Nvidia

August 12, 2026

Nvidia's Nemotron 4 aims for one trillion parameters, a scale Chinese labs already surpassed

Nvidia is working on Nemotron 4, a new open-weight model designed to rival the world’s best freely available models. The article Nvidia's Nemotron 4 aims for one trillion parameters, a scale Chinese labs already surpassed appeared first on The Decoder.

Read original article

Meta

August 12, 2026

Meta and Nvidia plant 'very firm flag' in open-weight AI race led by Chinese Labs - CNBC

Meta and Nvidia plant 'very firm flag' in open-weight AI race led by Chinese Labs CNBC

Read original article

OpenAI

August 12, 2026

Pakistani Judges Give Their Verdict on Judge GPT

Judges around the world have made headlines for illicitly using generative AI in their work. But in Pakistan, a large-scale trial of a specially designed AI tool for judges found the technology–together with appropriate training–boosted the number of cases resolved by 6.3% with no obvious drop in the quality of judgements.With a backlog of 2.26 million cases and fewer than two judges per 100,000 people – compared to 22 in the EU and 8 in Brazil – Pakistan’s judiciary was in sore need of help. So, in consultation with the judiciary, economist Sultan Mehmood, of the New Economic School in Moscow, Russia, and collaborators tested whether AI could ease the burden.They built a custom tool combining OpenAI’s GPT-4 large language model (LLM) with a knowledge base of nearly 130,000 Pakistani judicial opinions and statutes, to help judges with legal research and drafting judgements. They began offering the tool in 2024 to 1,559 trial judges – roughly half the country’s justices. “We do find an increase in cases resolved, and we don’t find any corresponding decrease in decision quality,” Mehmood says. First of its kind“It’s pretty amazing that he’s able to pull this off,” says David Autor, an economics professor at the Massachusetts Institute of Technology. “It’s not easy to do large-scale field experiments in civil service, but especially where the stakes are so high.” The 6.3% productivity boost is not overwhelming, he says, but it’s credible and likely to improve as the tool is more widely used.AI tools for judges are already being rolled out in Brazil and India, and prominent U.S. law professor Eric Posner has compared LLM judgments to human judgements in a single case study, but until now there have been no major independent assessment of ongoing judicial use of AI. The new study focused on Pakistan’s trial courts; Mehmood says judges there were enthusiastic from the start.“They were more techno-optimist than we were,” he says. “The delays are so huge this is something which they thought was worth trying anyway to reduce people’s suffering.”Some judges were also already using AI chatbots, Mehmood says, but commercial offerings performed poorly on Pakistani legal queries, frequently hallucinating case law. So the team built a tool tailored to the Pakistani context, called JudgeGPT.They used retrieval-augmented generation (RAG), which allowed the model to query a database of 128,292 Pakistani judicial opinions and 943 statutes. Responses included footnotes linking to cases and laws.“It turns out that actually the way to fix [hallucinations] isn’t just more intelligent models,” says study co-author Elliott Ash, an associate professor of law, economics, and data science at ETH Zurich in Switzerland. “It’s to attach the models to a tool that can do a search and verify the sources.” However, the researchers do not report hallucination rates.The team also put 1,197 judges through six 90-minute Zoom training sessions where Ash covered how LLMs work, their limitations, the risk of bias and hallucinations and the importance of verifying outputs. Another 180 judges only underwent general training on technology in legal research, while a final group got no training.By the time 487 judges had been through the program, the median district saw a jump of 6.3% resolved cases, and the more trained judges in a district, the bigger the effect. Appeal rates also fell slightly, suggesting faster resolution wasn’t leading to sloppier decisions. The JudgeGPT can be used to surface relevant case law with a simple text query and results provide links to the full text of the related judgements.Sultan Mehmood, Christoph Goessmann, and Elliott AshAddressing limitationsThe team also assessed the quality of judgements. Having legal experts evaluate large numbers of judgements was infeasible, Mehmood says, so the team asked OpenAI’s GPT-5-mini to choose between pairs of judgements from the same judge before and after training. The LLM chose post-training judgements 59% of the time. Two experienced Pakistani lawyers also evaluated the model’s analysis of 90 judgement pairs. They agreed with GPT-5-mini 70.6% of the time, compared to 73% agreement with each other.Training turned out to be vital. On average, JudgeGPT-trained judges logged in 56 times and sent 212 prompts over the study period, compared to 10 logins and 25 prompts after generic training. Those who had no training tended to use the tool for around a month and then drop off entirely, Mehmood says. “Just giving people the technology does not necessarily make them use it persistently,” he says.A 6.3% increase sounds modest, but the researchers calculated that a trained judge was resolving 38.5 more cases a month than the baseline, translating to roughly US $38.50 saved in judicial costs for every dollar spent running the tool. Ash also notes that these figures come from a nine-month period at the start of the trial, and that they’ve since updated both the underlying AI model and the database.For users the tool has been a lifeline. One participating trial judge, who spoke on condition of anonymity, says the number of cases assigned to them hasn’t dropped below 1,000 in more than a decade. The tool saves significant time, in particular searching for case law and summarising lengthy documents. “For research, it’s just one prompt away, whereas before I had to search for the precedents and laws for hours,” the judge says. “If I have to read 10 pages of a precedent, now I ask JudgeGPT to just summarise it for me and give me the crux, and it does that work in seconds.”But efficiency isn’t the only thing you want out of a justice system, says John Zeleznikow, professor of law and technology at La Trobe University in Australia. “What they’ve tried to do is be effective, [to] deal with more cases more quickly, and they’re able to do that,” he says. “What’s not that clear is whether what you call the quality of justice is better.”Zeleznikow saysAI can be useful but only if judges are diligent about evaluating and verifying the output. However, the working paper’s authors found that roughly a fifth of participant’sthe prompts given to JudgeGPT involved what they call “substantial AI delegation” – asking the tool what the best decision is, to produce legal reasoning or write opinions with little input from the judge. On the bright side, training lowered the proportion of inappropriate delegation.But given that judges are already using AI, Ash says better tools and training are crucial. “There are risks for using these AIs for sure, even with all these safeguards, but at some point you have to just put the judges in as strong a position as you can,” he says. “Have technological safeguards, but then try to encourage the judges not to rely on it too much.”

Read original article

Anthropic

August 12, 2026

I Asked AI to Write a Novel. It’s Not So Bad. - Mother Jones

I Asked AI to Write a Novel. It’s Not So Bad. Mother Jones

Read original article

OpenAI

August 12, 2026

Oh Lord, AI Reporters Are Actually Breaking Big News

Last week, an AI newsroom beat mainstream journalists—including WIRED—to a story about OpenAI and hacking. It’s just the beginning.

Read original article

Microsoft

August 12, 2026

Microsoft's new MAI Code 1.1 Flash gets crushed by Deepseek on both price and performance

Microsoft has released MAI Code 1.1 Flash, a code model for GitHub Copilot that's said to be 25 percent more token-efficient at a quarter of the cost of its predecessor. In benchmarks, though, it gets crushed by the cheaper Deepseek V4 Flash. The move fits a pattern: Microsoft talks up open AI, then bakes worse proprietary models into its apps to protect margins. The article Microsoft's new MAI Code 1.1 Flash gets crushed by Deepseek on both price and performance appeared first on The Decoder.

Read original article

Meta

August 12, 2026

Zuckerberg lays out Meta’s new vision for AI: Personal intelligence for all - Los Angeles Times

Zuckerberg lays out Meta’s new vision for AI: Personal intelligence for all Los Angeles Times

Read original article

Nvidia

August 12, 2026

NVIDIA AI Releases Nemotron 3.5 Lightning: A 30 B Open Mo E with 3 B Active Parameters, and Ne Mo Switchyard Model Router

NVIDIA's open 30B MoE targets the agent execution layer, with Switchyard routing each step to the cheapest capable model. The post NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router appeared first on MarkTechPost.

Read original article

OpenAI

August 12, 2026

From assistance to execution: How enterprises put AI to work

OpenAI research reveals how enterprises are adopting agentic AI, using ChatGPT and Codex, and how frontier firms are pulling ahead in AI adoption.

Read original article

Meta

August 12, 2026

SBCO: Self-Supervised, Verifier-Grounded Harness Optimization For Planning Agents

arXiv:2608.10157v1 Announce Type: new Abstract: Self-improving agents seek to reduce the human engineering effort behind AI systems by enabling them to evolve and self-improve their performance over time. Recently, methods like the Darwin G\"odel Machine and the Huxley G\"odel Machine have been proposed which enable open-ended, recursive self-improvement through self-reference where a coding agent edits its own code. Such self-referential self-improvement methods require that the competence required to perform the task coincides or aligns well with the competence required for self-modification which is the case for coding tasks. For domains or tasks, which do not satisfy the alignment needed, self-referential self-improvement is not available. In such cases, it is possible to adapt the above algorithms to other tasks by removing the self-referential aspect or introducing explicit self-modification of a meta-agent -- both computationally expensive, relying on population or self-modification search over many candidate agents. For planning tasks with explicit constraints, we propose a far cheaper alternative. We introduce SBCO (Self-supervised Block Coordinate Optimizer), a verifier-grounded harness optimizer in the same closed-loop, improve-from-experience family as the G\"odel-machine methods, but self-supervised rather than self-referential. Given an agentic harness, SBCO learns a decomposed bank of verifiers and a harness policy via approximate block coordinate ascent, improving the agent's outputs from its own graded feedback---with a fixed meta-agent and no human labels. Across two domains SBCO matches or exceeds a customized self-modifying baseline while using 4-5.5 times less compute budget.

Read original article

Claude

August 12, 2026

Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness

arXiv:2608.10008v1 Announce Type: new Abstract: LLM recommenders for top-$K$ item suggestion regularly emit titles outside the target catalog. Prior audits measure this as a binary out-of-domain rate; none ask whether the model knew it was hallucinating. We jointly audit hallucination rate (OOD@10) and verbalized-confidence calibration (ECE, Brier, reliability) for four zero-shot LLM recommenders from four independent vendors (Mistral Large, Llama-3.3-70B, GPT-OSS-120B, Claude Sonnet 4.6), not grounded or fine-tuned systems, across three catalogs (MovieLens-25M, Amazon Reviews 2023 Toys, Yelp Open Dataset), stratified by item popularity. Hallucination is catalog-dependent (0--0.2\% on MovieLens, 4.5--8.3\% on Amazon, 2.2--8.4\% on Yelp), but verbalized confidence is materially miscalibrated even when hallucination is zero (ECE up to 0.223 on MovieLens despite 0\% OOD). All four LLMs are systematically \emph{under}-confident across all twelve cells, verbalizing a mean of 67--86 on items they recommend with 92--100\% accuracy. This is the opposite of the over-confidence usually emphasized in LLM-hallucination work. The under-confidence is best read as an \emph{elicitation mismatch}: ``Just Ask'' elicits a generic recommendation-quality rating, not a catalog-membership probability. A conformal abstention threshold over verbalized confidence reduces hallucination by at most 0.7\,pp across $\alpha \in \{.05, .10, .15, .20\}$, at 4--21\,pp of coverage cost: the under-confident channel cannot separate correct items from hallucinations, so the threshold mostly removes correct items. We recommend that audits of LLM recommenders report calibration alongside OOD, and use catalog-anchored elicitation rather than generic confidence prompts.

Read original article

Meta

August 12, 2026

Connection Mind: Leveraging Social Networks and Large Language Models for Personalized Recommendation at Meta

arXiv:2608.10187v1 Announce Type: new Abstract: Modern recommendation systems on social media platforms such as Meta must model complex social relationships, including friendships, group memberships, and creator interactions, alongside massive and heterogeneous content such as text and video. Traditional recommendation models, however, often omit these signals or treat them independently, lacking the reasoning capability to integrate multi-relational context for fine-grained personalization. We present ConnectionMind, a production-ready recommendation framework that tightly integrates the social network structure with large language models (LLMs) to enable scalable, interpretable, and reasoning-aware personalization in Meta. ConnectionMind constructs a heterogeneous graph connecting users, items, friends, groups, and creator pages, and formulates recommendation as a graph reasoning problem: discovering personalized paths from users to candidate items. An LLM-based policy is employed to reason over these graph structures and guide recommendation decisions. To train the system at scale, ConnectionMind adopts a two-stage learning strategy. We first perform supervised fine-tuning (SFT) on large-scale user-item interaction trajectories to initialize the reasoning policy, followed by end-to-end reinforcement learning (RL) to refine the model's ability to reason over social graphs for personalized recommendation. Extensive experiments on multiple real-world datasets demonstrate the effectiveness of ConnectionMind compared to representative baselines. More importantly, ConnectionMind has been deployed in Meta's large-scale recommendation pipeline and has been evaluated through online A/B tests, achieving a 0.43% improvement in video watch time. These results demonstrate measurable real-world impact in a production recommendation system.

Read original article

Nvidia

August 12, 2026

Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation

arXiv:2608.10385v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as relevance assessors in information retrieval (IR) evaluation, raising questions about how assessor framing affects judgment reliability and downstream system comparison. We study persona conditioning as a diagnostic mechanism for exposing LLM assessor sensitivity. Using task-oriented personas drawn from two complementary sources (PersonaHub and NVIDIA Nemotron-Personas-USA), we instantiate five assessor roles emphasizing intent interpretation, domain expertise, contrastive judgment, evidence verification, and global search-quality assessment, compared with a standard UMBRELA baseline. Across six LLM backbones on TREC DL20 and RAG24, our analyses reveal structured rather than uniform assessor sensitivity. Judgments usually remain close to the baseline while shifting assessment strictness, evidential threshold, or interpretation emphasis rather than producing widespread relevance reversals. At the system level, high-capacity models preserve system-ranking agreement, while smaller models amplify persona-induced instability. Local rank-displacement analysis shows sensitivity concentrates on particular retrieval systems and system types, especially neural ranking/reranking systems on DL20 and RAG-oriented pipelines on RAG24. Persona source matters less than assessor role and model capacity. These findings position persona-conditioned judging as a controlled sensitivity probe for stress-testing LLM-based IR evaluation pipelines and identifying systems whose evaluation outcomes are sensitive to assessor framing.

Read original article

Meta

August 12, 2026

Artificial Intelligence in Gallbladder Imaging: A Rapid Evidence Review and Exploratory Meta-Analysis of Diagnostic Performance and Reader Assistance - Cureus

Artificial Intelligence in Gallbladder Imaging: A Rapid Evidence Review and Exploratory Meta-Analysis of Diagnostic Performance and Reader Assistance Cureus

Read original article

Nvidia

August 12, 2026

Nvidia’s Show of Financial Force Soothes Jittery Credit Markets - Bloomberg.com

Nvidia’s Show of Financial Force Soothes Jittery Credit Markets Bloomberg.com

Read original article

Nvidia

August 12, 2026

Day 0 Support for Qwen3.8-2.4 T-A95 B on v LLM

Day-0 vLLM support for Qwen3.8-2.4T-A95B: a 2.4-trillion-parameter hybrid MoE model served out of the box, with FP8/BF16 checkpoints plus NVFP4 and MXFP4 quantized weights, and co-developed kernels on NVIDIA and AMD hardware.

Read original article

Databricks

August 11, 2026

Taking AUTO CDC to the next level: Solving the hardest real-world use cases

Change data capture is one of the most common things data engineers build on Spark...

Read original article

Nvidia

August 11, 2026

Personalized AI startup River AI raises $1.1 B from consortium backed by Nvidia, AMD

River AI Inc., a startup that helps enterprises customize open-source artificial intelligence models, has raised $1.1 billion in early-stage funding. The company stated in today’s announcement that it received the capital over two rounds, a seed and a Series A. General Catalyst and AMP PBC were the lead investors. They were joined by Nvidia Corp., […] The post Personalized AI startup River AI raises $1.1B from consortium backed by Nvidia, AMD appeared first on SiliconANGLE.

Read original article

Google

August 11, 2026

Google’s Gemini AI app passes 1 billion monthly active users

Google LLC’s Gemini artificial intelligence app has passed 1 billion monthly active users, making it the 14th product in the company’s history to reach that mark. The company announced the milestone today in a blog post from Josh Woodward, vice president of Google Labs, Gemini and AI Studio. Chief Executive Sundar Pichai said in a […] The post Google’s Gemini AI app passes 1 billion monthly active users appeared first on SiliconANGLE.

Read original article

Meta

August 11, 2026

Your AI agent remembers everything. Here’s what happens when its owner changes.

Manus is becoming an independent company again after Chinese regulators ordered Meta to unwind its roughly $2 billion acquisition of The post Your AI agent remembers everything. Here’s what happens when its owner changes. appeared first on The New Stack.

Read original article

OpenAI

August 11, 2026

Accelerate cyber defense with Open AI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock | Artificial Intelligence - Amazon Web Services (AWS)

Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock | Artificial Intelligence Amazon Web Services (AWS)

Read original article