DataAIHub Daily

Archive →

August 31, 2026

50 curated AI news stories from leading AI companies.

Claude

August 31, 2026

Hackers Target Claude Accounts With Malware That Steals Login Sessions - PYMNTS.com

Hackers Target Claude Accounts With Malware That Steals Login Sessions PYMNTS.com

Read original article

Google

August 31, 2026

The Pentagon now has its own version of Chat GPT and Grok

Versions of OpenAI's ChatGPT and SpaceXAI's Grok will join Google's Gemini on the Pentagon's central portal for AI tools.

Read original article

OpenAI

August 31, 2026

Open AI Hugging Face Attack: 70,000 AI Agent Messages—‘Sacrifice Yes’

AI agents debated sacrifice, permadeath and collective goals while coordinating an attack—an extraordinary glimpse into AI agents coordinating in real time.

Read original article

Google

August 31, 2026

Google’s new forecasting model beats everyone. You can’t use it at work (yet).

On Monday, Google launched TimesFM-3, a 330-million-parameter time-series forecasting model trained on over a trillion real-world and synthetic data time The post Google’s new forecasting model beats everyone. You can’t use it at work (yet). appeared first on The New Stack.

Read original article

Anthropic

August 31, 2026

The Music Industry's New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With Fear - Futurism

The Music Industry's New Lawsuit Against Anthropic Should Have Dario Amodei Shivering With Fear Futurism

Read original article

Nvidia

August 31, 2026

Space X designed an orbital Vera Rubin. Radiation comes next.

SpaceX and Nvidia say they are adapting the Vera Rubin NVL72 rack-scale AI platform for orbital use, with SpaceX targeting The post SpaceX designed an orbital Vera Rubin. Radiation comes next. appeared first on The New Stack.

Read original article

OpenAI

August 31, 2026

Open AI wants to charge only when AI gets it right — here’s the catch

AI companies have always charged customers for the tokens they use, whether the model gives them exactly what they need The post OpenAI wants to charge only when AI gets it right — here’s the catch appeared first on The New Stack.

Read original article

Anthropic

August 31, 2026

Anthropic IPO Could Open AI Listings Floodgates: Madrona

The IA40 list from Madrona is back, spotlighting the private companies shaping the next phase of AI. Matt McIlwain, managing director at Madrona, explains how the list is assembled, why OpenAI, Anthropic and Databricks dominate AI fundraising, and how a successful Anthropic IPO could open the door for a new wave of public offerings in 2027. He joins Ed Ludlow on "Bloomberg Tech." (Source: Bloomberg)

Read original article

Anthropic

August 31, 2026

Nvidia Deepens Chip Ties With $3.5 Billion Media Tek Bet | Bloomberg Tech 8/31/2026

Bloomberg’s Ed Ludlow breaks down Nvidia's $3.5 billion investment in MediaTek, deepening its collaboration with the Taiwanese chipmaker and demonstrating how the AI leader plans to tackle Big Tech efforts to build custom silicon. Plus, a look at Tim Cook's last day as CEO of Apple before John Ternus takes the reins on September 1st. And, all eyes on Anthropic as it prepares to file publicly in the coming weeks. (Source: Bloomberg)

Read original article

Anthropic

August 31, 2026

“Zlibrary my beloved”: Anthropic staff chats extolling piracy cited in Sony suit

Lawsuit: Anthropic’s torrenting totally screwed songwriters as AI songs top charts.

Read original article

Google

August 31, 2026

Hurricane model developed by Google boosted with artificial intelligence showing promise - WGCU

Hurricane model developed by Google boosted with artificial intelligence showing promise WGCU

Read original article

OpenAI

August 31, 2026

Hugging Face hack could indicate cultural issues at Open AI

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. By now you’ve probably heard about last month’s major AI security incident, in which OpenAI agents escaped their sandbox and hacked into the AI platform Hugging Face while trying to cheat on…

Read original article

OpenAI

August 31, 2026

Open AI Lets Some Customers Pay Only When AI Performs - PYMNTS.com

OpenAI Lets Some Customers Pay Only When AI Performs PYMNTS.com

Read original article

Anthropic

August 31, 2026

China blasts Anthropic and sets tough terms for U.S.-China AI safety talks - Los Angeles Times

China blasts Anthropic and sets tough terms for U.S.-China AI safety talks Los Angeles Times

Read original article

Anthropic

August 31, 2026

Forrester CEO Sizes Up the AI Boom

Nvidia makes another big bet on the future of AI with a $3.5 billion investment in MediaTek, as investors gear up for a potentially record-breaking Anthropic IPO. Forrester Research Founder and CEO George Colony weighs in on the AI boom, whether the hype is justified and what comes next. He speaks on Bloomberg Open Interest. (Source: Bloomberg)

Read original article

OpenAI

August 31, 2026

Open AI Says Ad Business Reaches $1 Billion Run Rate - PYMNTS.com

OpenAI Says Ad Business Reaches $1 Billion Run Rate PYMNTS.com

Read original article

Meta

August 31, 2026

Instagram admits users often can't tell AI profiles from real people

Instagram is replacing its "AI creator" tag with a new "AI-generated profile" label because users can't tell AI profiles from real people. Singularity, defined by Instagram user competence. Profiles without the label get their reach and recommendations throttled. As recently as late 2024, Meta was still planning a coexistence of AI characters and humans on the platform. The article Instagram admits users often can't tell AI profiles from real people appeared first on The Decoder.

Read original article

Databricks

August 31, 2026

Autoscaling Lakebase Postgres

Choosing a database instance size before you know the workload is an old building pattern...

Read original article

Nvidia

August 31, 2026

An Exclusive Interview with Jensen Huang on Nvidia’s AI Future | Open Interest 8/31/2026

Get a jump start on the US trading day with Dani Burger on "Bloomberg Open Interest." Bond investors aren’t convinced Fed Chair Kevin Warsh is ready to raise rates. Nvidia doubles down on AI, deepening ties with MediaTek — we speak exclusively with CEOs Jensen Huang and Rick Tsai. Plus, Newell Brands on cutting marketing costs 80% with AI, and Standard Lithium on the race to build America’s lithium supply. (Source: Bloomberg)

Read original article

Anthropic

August 31, 2026

Anthropic’s Mega-IPO Plan Looms Over Packed US Listing Calendar

Anthropic PBC’s IPO is casting a long shadow over companies’ US listing plans, as they try to find room for their deals to grab attention after the Sept. 7 Labor Day holiday. The Claude developer is preparing to file publicly for an initial public offering that’s expected to raise as much as SpaceX’s record $86.2 billion debut, if not more, in the coming weeks. Some firms and their backers are finding it hard to get the attention of long-term-oriented investors and sovereign wealth funds when the prospect of an Anthropic IPO is imminent, according to people familiar with the preparations. For a closer look, we speak with Bailey Lipschultz, Senior Equities Reporter for Bloomberg News. (Source: Bloomberg)

Read original article

Nvidia

August 31, 2026

Run NVIDIA Bio Ne Mo NIM Microservices for Protein Structure Prediction in Claude Science

Agentic AI is changing how research is done. AI scientists can read papers, propose hypotheses, call models, and determine which experiments to prioritize next....

Read original article

Nvidia

August 31, 2026

The AI Chip War's New Front: Control The Cloud, Not The Silicon

A draft U.S. rule would block China from renting Nvidia GPU compute via data centers in Thailand and Singapore. The AI chip war's next front: the cloud, not the silicon.

Read original article

Nvidia

August 31, 2026

Scale AV Perception Across Vehicle Platforms with NVIDIA Omniverse Nu Rec

A perception stack is shaped by the vehicle that carries it. Move the same software to a new carline—for example, from an SUV to a sedan or another vehicle...

Read original article

Google

August 31, 2026

Cloud CISO Perspectives: Tips on securing the water sector in the AI era

Welcome to the second Cloud CISO Perspectives for August 2026. Today, Chris Sistrunk and Stephanie Kiel detail the critical issues facing the water sector, and actionable steps that OT operators can take to secure their infrastructure.As with all Cloud CISO Perspectives, the contents of this newsletter are posted to the Google Cloud blog. If you’re reading this on the website and you’d like to receive the email version, you can subscribe here. aside_block <ListValue: [StructValue([('title', 'Get vital board insights with Google Cloud'), ('body', <wagtail.rich_text.RichText object at 0x7f624230f1f0>), ('btn_text', 'Visit the hub'), ('href', 'https://cloud.google.com/solutions/security/board-of-directors?utm_source=cgc-site&utm_medium=et&utm_campaign=FY26-Q2-GLOBAL-GCP39634-email-dl-dgcsm-CISOP-NL-177159&utm_content=-&utm_term=-'), ('image', <GAEImage: GCAT-replacement-logo-A>)])]> Tips on securing the water sector in the AI eraBy Chris Sistrunk, Practice Leader, OT, Mandiant Consulting, and Stephanie Kiel, Head of Cloud Security Policy, Government Affairs and Public Policy, Google Cloud Chris Sistrunk, Practice Leader, OT, Mandiant Consulting Google Cloud’s threat intelligence teams have observed that threat actors are becoming bolder when targeting critical infrastructure amid geopolitical conflicts. Recently, we’ve seen increased targeting of water utilities' internet-connected programmable logic controllers in the U.S. Stephanie Kiel, Head of Cloud Security Policy, Government Affairs and Public Policy, Google Cloud Historically, cyber incidents haven’t usually disrupted operations, in part because water utility operators have long had manual override capabilities and established water-quality checks that kick in before water reaches consumers. Pumps and pipes fail routinely for reasons that have nothing to do with cyber threats. However, they do require our urgent attention and a commitment to stronger security hygiene. Manual overrides provide a reliable safety net, but preventing cyber threats still requires a commitment to fundamental digital security — especially in the AI era. We recommend a threat-informed, risk-managed response. The current state of water sector security is indicative that additional action should be strongly considered in light of the unique operational resilience that keeps these systems safe. Actions water and wastewater utilities should consider For resource-constrained utilities, the most effective defense is to focus on cybersecurity fundamentals. By prioritizing these fundamental practices, you can significantly harden your systems and transform your organization into a far more challenging and resilient target, causing even well-resourced threat actors to look elsewhere. Inventory assets and assess exposure: Identify if your control systems are insecurely exposed to the internet, which often allows for the successful exploitation of vulnerabilities. Basic security hygiene: Replace default credentials with strong passwords, and rigorously harden exposed access points, including firewalls. Backups: Make sure that critical systems, including control systems, are safeguarded following the proven 3-2-1 backup rule (keep three copies of your data on two types of storage, with at least one copy stored off-site). Ensure critical spare equipment is on-hand to minimize downtime from cyberattacks. Segmentation: Use network segmentation and multifactor authentication to ensure that remote access, when necessary, is strictly controlled. You should use read-only access where full control isn't required. Emergency planning: Integrate cyber-incident planning into your existing all-hazards incident command system, including FEMA NIMS and Incident Command System for Industrial Control Systems, the same response structures you already use for physical pipe breaks, boil water alerts, and natural disasters. Secure third-party and vendor access: As many water utilities do not manage their own IT or OT and rely on third-party system integrators, you should audit the remote connections used by the system integrators and maintenance contractors. You should ensure third-party vendors are held to rigorous access controls (such as MFA standards) and logging requirements. These recommendations echo guidance from the American Water Works Association, the National Rural Water Association, the Water-ISAC, the Environmental Protection Agency, the Cybersecurity and Infrastructure Security Agency, and the FBI. Recommendations for IT and OT leaders: Bridging the governance gap IT and OT leaders must work together to build a unified governance framework and should focus on making cyber-physical systems more resilient over the long term, a collective effort that spans government agencies, private sector organizations, and individuals. The goal is to build a future where these systems are secure, adaptable, and capable of recovering quickly from disruptions. Although PLCs almost always sit outside standard software development practices, a robust approach to the software your organization uses can significantly enhance your overall security posture, such as those outlined in NIST’s Secure Software Development Framework (SSDF). They’re also good examples of leading indicators that can help you gauge your resilience, and to help you get started we’ve published a guide to evaluate leading indicators. Manual overrides provide a reliable safety net, but preventing cyber threats still requires a commitment to fundamental digital security — especially in the AI era. As technology evolves, it is critical to modernize security, transitioning from a reactive, manual model to an AI-augmented approach that keeps human expertise central to decision-making. This approach offers an unique opportunity to be a force multiplier for lean security teams. To stay ahead of today’s threats, organizations must move beyond simple compliance checklists and adopt a more agile, threat-informed strategy that makes compliance a natural outcome of good security, rather than the primary goal. The Mandiant Operational Technology (OT) Theory of 99 has become more relevant in the AI era. Although the funnel of opportunity has been significantly compressed, in intrusions that go deep enough to impact OT: 99% of compromised systems will be computer workstations and servers 99% of malware will be designed for computer workstations and servers 99% of forensics will be performed on computer workstations and servers 99% of detection opportunities will be for activity connected to computer workstations and servers 99% of intrusion dwell time happens in commercial, off-the-shelf computer equipment before any Purdue level 0-1 devices are impacted As a result, there is often a significant overlap across tactics, techniques, and procedures used by threat actors who target IT and OT networks. However, the Theory of 99 underscores a significant defender's advantage in the AI era. By using advanced AI capabilities to secure the 99% of intermediary infrastructure, organizations can proactively neutralize threats and ensure robust protection for the critical 1% of physical operational processes. AI for cyber defense As we have shared before, AI capabilities offer the opportunity to shift the balance in network security in the favor of defenders. The defender’s advantage becomes even more important as malicious actors increasingly use AI capabilities across the attack lifecycle. In the current threat environment, automating defenses can serve as a force multiplier for human security teams, enhancing decision-making and productivity to ensure critical exposures are addressed before they can be exploited. With careful planning, critical infrastructure providers can protect their physical assets while building a more resilient, threat-informed defense. To effectively realize AI advantages for defense, you should integrate AI tools into systems in a structured, intentional way. It’s crucial that operators understand the unique vulnerabilities that AI introduces to physical processes, evaluate specific business uses that can benefit from security automation, and establish clear frameworks to continuously test and monitor. As part of our approach, we’ve developed the Secure AI Framework to help you achieve secure integration and deployment of AI capabilities, regardless of sector. Most importantly, human oversight must remain central — meaning that AI should support decision-making, and safety practices need to be embedded directly into incident response plans. What’s next for water security Protecting water systems from malicious cyber threats is not just a technical challenge; it is a fundamental public safety imperative. Given that access to clean, reliable water is an essential service, we anticipate that federal, state, and local governments will increasingly shift from policy debate to decisive action to ensure the continuity of this critical public infrastructure in the face of cyber threats. For example, the Office of the National Cyber Director in partnership with the State of Texas has just launched a pilot program to help protect water infrastructure providers from cyberattacks, and U.S. senators have already introduced a new bill in response to recent events. Google is committed to helping you protect your cloud and hybrid cloud OT environments. To learn more about Google guidance on securing critical infrastructure, please visit our CISO Insights Hub. aside_block <ListValue: [StructValue([('title', 'Learn something new'), ('body', <wagtail.rich_text.RichText object at 0x7f624230f310>), ('btn_text', 'Watch now'), ('href', 'https://x.com/googlecloud/status/2090213589558698309?s=20'), ('image', <GAEImage: Cloud-CISO-Perspectives-logo-A>)])]> In case you missed itHere are the latest updates, products, services, and resources from our security teams so far this month:Empowering autonomous agents with advanced security governance: To be useful and secure, AI agents need access — and also guardrails. In our new State of AI infrastructure report, 79% of tech leaders cite security, governance, or operations as their most significant challenge to scaling inference. Read more.The state of cloud risk 2026: Most security findings aren’t real attacker opportunities: Wiz Research telemetry reveals why the majority of high-severity findings lack a path to compromise. Read more.Introducing Google Cloud Fault Injection Testing in preview: When databases fail and network paths falter, you still need your mission-critical cloud services to stay online. Fault Injection Testing (FIT) can help you automate failure testing to ensure predictable behavior during disruptions. Read more.How Wiz built AI-powered data discovery: Inside the multi-agent pipeline and feedback loops that turned a bucket scanner into a context engine. Read more.Democratizing FinOps with Wiz: How the Wiz Cloud Cost automates cost allocation to power developer-led cost optimization and connect cost to business value. Read more.Defend against agent risks with layered protections in Google Workspace Studio: Studio incorporates layered defenses to mitigate risks from threat actors and robust observability tools to help organizations adopt agents safely. Built on Google’s secure-by-design architecture, Studio combines native threat defenses with deep ecosystem visibility to secure multi-step agentic workflows. Read more.Please visit the Google Cloud blog for more security stories published this month. aside_block <ListValue: [StructValue([('title', 'Join the Google Cloud CISO Community'), ('body', <wagtail.rich_text.RichText object at 0x7f624230f220>), ('btn_text', 'Learn more'), ('href', 'https://rsvp.withgoogle.com/events/google-cloud-ciso-community-interest-form-2026?utm_source=cgc-blog&utm_medium=blog&utm_campaign=FY25-Q1-global-GCP30328-physicalevent-er-dgcsm-parent-CISO-community-2025&utm_content=cisop_&utm_term=-'), ('image', <GAEImage: GCAT-replacement-logo-A>)])]> Threat Intelligence newsDistinct clusters target individuals of interest to Russia: Google Threat Intelligence Group (GTIG) is tracking three suspected Russian cyber espionage threat clusters abusing legitimate authentication flows to target individuals working in academia, aerospace, governments, and think tanks across Europe and in the U.S. Read more.Inside 90 days of attacks on AI infrastructure: Wiz honeypots uncover active campaigns targeting LiteLLM, MCP servers, and AI frameworks through RCE, blind prompt injection, and memory credential theft. Read more.Version Control DFIR: A cheatsheet to GitHub, GitLab, Bitbucket, and Azure DevOps: A practitioner’s guide to log visibility, incident readiness, and threat hunting across the major version control services. Read more.Rust supply chain attack on arrayref: Significant overlap with DPRK campaigns: Malicious versions of the arrayref Rust crate (and others) executed a backdoor at compile time. The campaign's infrastructure overlaps with recent DPRK supply chain attacks, including Mastra and axios. Read more.Please visit the Google Cloud blog for more threat intelligence stories published this month. Now hear this: Podcasts from Google CloudCloud Security Podcast: Patching browsers with AI, agents, Rust, and your tabs: Jasika Bawa and Doug Turner of Chrome Security explore how Google Chrome now uses AI agents to autonomously identify and patch security vulnerabilities at an unprecedented scale, significantly accelerating the browser's update cadence. Listen here.Cloud Security Podcast: All about Project Atlas, Wiz's AI vulnerability research: Near Orfeld, head of vulnerability research, Wiz, discusses how his team uses multi-agent AI systems for discovering high-impact zero-day vulnerabilities in cloud infrastructure. Listen here.Cloud Security Podcast: How Google eliminates classes of vulnerabilities at scale: How do you build the foundations for a secure Google-scale enterprise that stays secure even if an AI is writing the code and nobody has time to review it? Christoph Kern, principal security engineer, Google, explores what secure-by-design really means in the AI era. Listen here.To have our Cloud CISO Perspectives post delivered twice a month to your inbox, sign up for our newsletter. We’ll be back in a few weeks with more security-related updates from Google Cloud.

Read original article

Google

August 31, 2026

From weeks to minutes: The new agentic era of data pipelines

Data pipelines are the backbone of the modern enterprise, yet a barrier to entry exists for orchestrating them, making this critical capability unavailable to many data professionals. Following our announcements at Google Cloud NEXT ’26, where we introduced the Orchestration Pipelines framework, we are fundamentally changing this dynamic. To bring this powerful framework directly to practitioners, we offer the Data Agent Kit — a unified, freely available, and open-source collection of data engineering and data science tools that integrate directly into your preferred IDE or CLI (such as VS Code, Claude Code, or Codex). The Data Agent Kit seamlessly embeds the Orchestration Pipelines framework into your workflow in two distinct ways. First, it provides a dedicated Data Engineering tab for comprehensive pipeline management. Second, it includes a specialized agentic skill designed to author, deploy, and troubleshoot production-grade Apache Airflow® DAGs using natural language. By pairing these specialized agent skills with a declarative YAML DSL, all data personas — from analysts to ML engineers — can bypass complex Python Airflow boilerplate. This framework decouples high-level orchestration logic from underlying compute execution, democratizing access to powerful MLOps capabilities across your entire data organization. In this post, we will walk through an exemplary MLOps use case to demonstrate how easily this can be achieved. Setting up your environment Before authoring your first Orchestration Pipeline, you need to set up your local development environment. Getting started takes less than two minutes. 1. Install and configure the extension To install the extension in your preferred IDE or CLI — such as VS Code, VS Code forks, Antigravity, Claude Code, Antigravity CLI, or Codex — and authenticate it with your Google Cloud account, follow the step-by-step setup guide in the official documentation: Google Cloud Data Agent Kit installation guide 2. Verify orchestration pipeline skills Once installed, verify that the required agent skills are active: Open the ‘Google Cloud Data Agent Kit’ panel on the VS Code activity bar. Navigate to ‘Settings’ then ‘Skills’. Ensure the ‘gcp-pipelines-orchestration’ skill is enabled. This skill provides the agent with deep contextual knowledge of pipeline syntax, variable substitution, secret management, and automated incident diagnosis for Airflow runs. 3. Building your first pipeline To start authoring, building, and validating orchestration pipelines directly inside the any VS Code compatible IDE using natural language prompts, follow the official building guide: Build pipelines guide An example business problem: Proactive supply chain management Let’s walk through an example business problem. In the logistics and retail sector, customer satisfaction hinges on accurate delivery estimates. When an order is delayed without warning, customer churn can spike and support costs can escalate. To address this, we are building an end-to-end MLOps architecture that predicts the exact transit time (in days) based on warehouse location, customer location, and order characteristics. By predicting these delays before shipping, operations teams can proactively notify customers or automatically upgrade shipping tiers before Service Level Agreements (SLAs) are breached. To make this architecture fully reproducible, we use the bigquery-public-data.thelook_ecommerce public dataset in BigQuery. For demo purposes, we split this static dataset into training and inference sets. In a real-life scenario, inference would be performed on new, incoming data. This dataset provides authentic operational complexity: Geographical data: Latitude and longitude for both customer addresses (users) and distribution centers (distribution_centers). Temporal data: Granular order lifecycle timestamps (created_at, shipped_at, delivered_at). Order attributes: Product categories, pricing, and fulfillment status (orders, order_items). By combining this dataset with BigQuery, Managed Service for Apache Spark serverless, Gemini Enterprise Agent Platform, and dbt, we will demonstrate how to build an automated, self-healing MLOps loop that handles training, daily batch inference, and model drift evaluation. The agentic workflow: From prompt to pipeline in minutes With the extension configured, we can bypass boilerplate Python for DAG authoring entirely. Inside VS Code, we opened the Data Agent Kit chat and provided a single natural language prompt to define our continuous MLOps feedback loop: Note: The detailed prompt was crafted with repeatability in mind specifically for this blog post. In real-life scenarios, you can achieve the same result in a more conversational way, pipeline by pipeline. The complete prompt and all generated files are available in the Orchestration-pipelines GitHub repository. Note: While frontier models equipped with the Orchestration Pipelines skill can often scaffold complete workflows in a single step, LLM responses naturally vary based on model versions, workspace context, and token depth. If a specific parameter, dataset path, or dependency is omitted in the initial pass, simply provide a short follow-up prompt. Within minutes, the Data Agent Kit generated the underlying PySpark scripts, dbt configurations, and the three declarative YAML pipelines. Please find below the generated YAML pipelines and a visual diagram of them. This pipeline is a simplified example designed to showcase Orchestration Pipelines capabilities. In practice, recommended production MLOps setups will vary depending on your specific use cases and operational needs. Pipeline 1: The training engineThis pipeline serves as our heavy-compute engine. The agent generated a YAML definition that first queries BigQuery to extract historical completed orders. It then dynamically provisions a Managed Spark serverless cluster to calculate geographical distances and train a model for production use. Finally, it pushes the trained model to Gemini Enterprise Agent Platform Model Registry. code_block <ListValue: [StructValue([('code', 'modelVersion: "1.0"\r\npipelineId: "training-pipeline"\r\nrunner: airflow\r\nowner: "mlops"\r\ntags:\r\n - "job:datacloud:antigravity"\r\ndefaults:\r\n projectId: "your-project-id"\r\n location: "us-central1"\r\n executionConfig:\r\n retries: 0\r\n\r\nactions:\r\n - sql:\r\n name: "extract_training_data"\r\n engine:\r\n bigquery:\r\n location: "US"\r\n destinationTable: "your-project-id.mlops.training_dataset"\r\n query:\r\n path: "blogpostdemo/training_query.sql"\r\n\r\n - pyspark:\r\n name: "train_model_dataproc"\r\n dependsOn:\r\n - "extract_training_data"\r\n engine:\r\n dataprocServerless:\r\n location: "us-central1"\r\n resourceProfile:\r\n inline:\r\n runtimeConfig:\r\n version: "2.3"\r\n properties:\r\n "spark.dataproc.driverEnv.PYTHONPATH": "./libs/lib/python3.11/site-packages"\r\n "spark.executorEnv.PYTHONPATH": "./libs/lib/python3.11/site-packages"\r\n mainFilePath: "blogpostdemo/train_model.py"\r\n environment:\r\n requirements:\r\n inline:\r\n list:\r\n - "tensorflow==2.14.1"\r\n - "numpy<2.0.0"\r\n - "protobuf<5.0.0dev"\r\n - "google-cloud-storage"\r\n\r\n - ai:\r\n name: "upload_model_vertex"\r\n dependsOn:\r\n - "train_model_dataproc"\r\n agentPlatform:\r\n projectId: "your-project-id"\r\n location: "us-central1"\r\n modelUpload:\r\n modelName: "transit_days_predictor"\r\n modelArtifactUri: "gs://your-bucket-name/models/tf_transit_days_model"\r\n : "us-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-14:latest"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f624223feb0>)])]> Pipeline 2: Daily inferenceFor our daily operational workflow, this lightweight pipeline applies the trained model to all currently in-transit orders. It queries the dataset via BigQuery job, executes inference job via Gemini Enterprise Agent Platform, and writes the results back to a BigQuery table to flag potential SLA breaches for the customer support team. code_block <ListValue: [StructValue([('code', 'modelVersion: "1.0"\r\npipelineId: "inference-pipeline"\r\nrunner: airflow\r\nowner: "mlops"\r\ntags:\r\n - "job:datacloud:antigravity"\r\ndefaults:\r\n projectId: "your-project-id"\r\n location: "us-central1"\r\n executionConfig:\r\n retries: 0\r\n\r\nactions:\r\n - sql:\r\n name: "extract_inference_data"\r\n engine:\r\n bigquery:\r\n location: "US"\r\n destinationTable: "your-project-id.mlops.inference_dataset"\r\n query:\r\n path: "blogpostdemo/inference_query.sql"\r\n\r\n - ai:\r\n name: "run_vertex_batch_prediction"\r\n dependsOn:\r\n - "extract_inference_data"\r\n agentPlatform:\r\n projectId: "your-project-id"\r\n location: "us-central1"\r\n batchInference:\r\n jobDisplayName: "inference_job"\r\n modelName: "projects/your-project-id/locations/us-central1/models/your-model-id"\r\n bigquerySource: "bq://your-project-id.mlops.inference_dataset"\r\n : "bq://your-project-id.mlops"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f624223f550>)])]> Pipeline 3: Automated evaluation and branchingThe daily evaluation pipeline acts as our automated quality gate. It triggers dbt models to join our predictions with actual delivery timestamps, calculating absolute errors and SLA breaches. Using built-in logic, the pipeline automatically evaluates these metrics. If the model’s error rate exceeds our acceptable threshold, it conditionally triggers the ‘training-pipeline’ to generate a fresh model. code_block <ListValue: [StructValue([('code', 'modelVersion: "1.0"\r\npipelineId: "evaluation-pipeline"\r\nrunner: airflow\r\nowner: "mlops"\r\ntags:\r\n - "job:datacloud:antigravity"\r\ndefaults:\r\n projectId: "your-project-id"\r\n location: "us-central1"\r\n executionConfig:\r\n retries: 0\r\n\r\nactions:\r\n - pipeline:\r\n name: "run_dbt_models"\r\n framework:\r\n dbt:\r\n airflowWorker:\r\n : "blogpostdemo/dbt_project"\r\n\r\n - python:\r\n name: "check_retraining_condition"\r\n dependsOn:\r\n - "run_dbt_models"\r\n mainFilePath: "blogpostdemo/evaluate_drift.py"\r\n pythonCallable: "check_drift"\r\n engine:\r\n local: {}\r\n\r\n - :\r\n name: "trigger_retraining_pipeline"\r\n dependsOn:\r\n - "check_retraining_condition"\r\n pipelineId: "training-pipeline"\r\n bundleId: "my-first-bundle"\r\n waitForCompletion: false'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f624223f3a0>)])]> Automated deployment to Managed Service for Apache Airflow Authoring pipeline logic is only half the battle; deploying it securely and reliably to production is where data teams historically lose valuable time. With Orchestration Pipelines, deployment is streamlined through standard CI/CD practices. Rather than manually writing deployment scripts or configuring complex environment boundaries, the Data Agent Kit automatically generates the necessary continuous integration workflows (such as GitHub Actions) for your workspace. This means you can simply click commit, and the framework will seamlessly package and deploy your Orchestration Pipeline bundle directly to your Managed Airflow environment. For a comprehensive guide on integrating these automated workflows into your existing CI/CD pipelines, review the official guide: Deploying Orchestration Pipelines. Day-two operations: Monitoring and agentic troubleshooting Maintaining these pipelines is just as intuitive as building them. By bringing the orchestration control plane directly into your IDE, the Data Agent Kit provides real-time monitoring of your Managed Airflow runs without requiring you to constantly context-switch between browser tabs. The Data Agent Kit provides real-time monitoring of your Managed Airflow runs directly within your IDE. The Data Agent Kit visualises the created pipeline. Inevitably, infrastructure or data issues occur—perhaps a Managed Spark cluster hits an out-of-memory exception due to a seasonal data spike, or a BigQuery quota is reached. Resolving these issues no longer requires digging through thousands of lines of raw execution logs. If a pipeline fails, the Data Agent Kit provides out-of-the-box agentic troubleshooting. With the click of a "Troubleshoot" button in your IDE, the Data Engineering Agent analyzes the failure context. It can accurately distinguish between infrastructure quota limits and code-level bugs, instantly providing a root-cause summary and suggesting an inline fix (such as scaling up the compute template). Agentic troubleshooting instantly diagnoses pipeline failures, identifies infrastructure bottlenecks, and suggests inline fixes. Summary: Accelerating time to value Building a resilient MLOps architecture — extracting historical data, executing dbt transformations, provisioning Managed Spark ML compute, integrating Gemini Enterprise Agent Platform for model registry and inference, and configuring cross-DAG conditional triggers — traditionally takes platform engineering teams weeks of writing complex Python Operator logic. With Orchestration Pipelines and the Data Agent Kit, this entire lifecycle was authored, deployed, and easily maintained in a matter of minutes. By replacing boilerplate infrastructure code with a declarative, agent-ready standard, we are ensuring your data organization spends less time orchestrating pipelines and more time delivering tangible business value. Get Started Today: Review the Orchestration Pipelines documentation. Install the Data Agent Kit in your preferred IDE or CLI and configure your workspace. Learn more about the broader ecosystem in our recent blog post: Data Agent Kit brings data skills and tools to your IDE or CLI. Explore reference architectures in the Data Agent Kit documentation.

Read original article

Google

August 31, 2026

What’s new in AI infrastructure and orchestration in August

Welcome back to What’s new in AI infrastructure and orchestration this month, a collection of product updates, how-tos, customer stories, research and other resources about all the AI compute, networks, storage, frameworks, and orchestration software that you can find at Google Cloud. To be honest, we thought August would be a slow month, but nothing could be further from the truth. Read on and you’ll see what we mean. August 2026 Product, technology, and tools updates Product update: Filestore, Google Cloud’s first-party, secure, scalable NFS file service, has emerged as a popular storage platform for AI and agentic workflows, and now, it’s even better suited to the task, with a new backend storage layer built directly on Colossus, Google’s foundational distributed storage system. This new backend lets you provision IOPS independently from storage capacity, and is deeply integrated with GKE. In AI environments, this can help you service so-called agentic swarms — large groups of agents that need to read and write to a common dataset — without a drop off in performance. For more, check out the blog post. New feature: gVisor sandboxes are now available in distributed Ray clusters on GKE. In partnership with Anyscale, we introduced an experimental library for Ray that brings gVisor, Google’s open-source application kernel, directly into distributed Ray clusters. gVisor provides lightweight environments with stronger isolation than ordinary containers, plus fast startup times and low memory overhead. To try out these sandboxing capabilities on GKE, head over to the Ray sandboxing User Guide. Product update: Looking for high-performance, easy-to-use infrastructure on which to run a personal AI agent, but don’t want to spend a lot of money? New Cloud Run instances are dedicated, singleton compute runtimes on Cloud Run that won’t shut down when the agent is idle. Better yet, the cost to run a Cloud Run instance with 1 vCPU and 1 GiB of memory continuously for 30 days is just $5.70. Practitioner guides, documentation and how-tos How-to guide: Big news in Model Context Protocol (MCP) land: As of the 2026-07-28 specification, the protocol core is “completely stateless. The handshake is gone. The initialize / initialized handshake (SEP-2575) and the logical Mcp-Session-Id header (SEP-2567) have been removed entirely. Instead, every request is now self-describing and independent.” Whoa. Learn more about the changes that the latest MCP specification brings, and more importantly, how to implement them, in this Google Developers blog. Guide: Real-time AI systems make a mess of traditional network load balancing techniques. “Instead of handling isolated requests, the backend has to manage a continuous, live bidirectional stream. You’re dealing with a constant stream of audio chunks, transcripts, model outputs, and synthesized speech flowing back and forth simultaneously.” Things only get worse when the user gets involved. “The server has to immediately halt its current speech generation, pivot to update the context, maybe trigger a new tool, and start drafting a different response; this must be done without dropping the connection.” For a new approach to managing load in the AI era, read Scaling real-time AI agents with session-aware load balancing. How-to: Learn how to build an elastic, scalable LLM inference platform on GKE, even with a mix of different GPU accelerators. The proposed architecture combines Capacity Advisor and Compute Advisor, plus high-performance storage like RunAI:model streamer or GCPFuse with parallel downloads. Get all the details here. Documentation: The thing about hosts with GPUs or TPUs is that you can’t use live migration to update them, setting up a maintenance challenge. In this new docs page, learn how to update accelerator-equipped hosts according to your tolerance for downtime for your training and inference workloads. Documentation: Advanced Compute Images, or ACIs, are standardized image stacks for AI/ML and HPC infrastructure, so you don’t need to manually build your own custom images. In this new docs page, learn how to create an ACI image using the Google Cloud CLI, console, or SchedMD's Slurm workload manager. Guide: AI workloads are notoriously difficult to architect, resource-intensive, and bursty, which can also lead to scaling bottlenecks and large pools of underutilized — or misutilized — compute resources. A new blog outlines the three main ways to achieve dynamic capacity management in Google Cloud: 1) scheduling capacity for planned downtime; 2) maintaining automated fallback capacity for unplanned downtime; and 3) relying on GKE’s core orchestration capabilities to automate resource allocation. Customer and partner updates Business orchestration software provider UiPath was dealing with spiky workloads, and wanted more predictable costs. To get there, it re-architected its infrastructure, moving from isolated clusters to a shared Google Cloud GPU fleet that included both A3 VM instances (NVIDIA H100 GPUs) for training with G4 VM instances (NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs) for inference. You can read more about their architecture here. Mirendil, an frontier AI lab focused on accelerating AI development, announced that it is using AI Hypercomputer with both TPUs and NVIDIA GPUs to support its model pre-training and post-training applications. Replenit, a retail CRM provider, built its AI decision engine in Google Cloud, using BigQuery, Gemini Enterprise Agent Platform, and open-source Gemma models that it runs on Cloud TPUs. This latter combination provided Replenit with 90% lower pipeline costs than their previous cloud provider, the company reports. Read the full case study for more. Malachyte architected its AI-powered e-commerce recommendation platform on top of Bigtable, Managed Service for Apache Kafka, Pub/Sub, Compute Engine, and last but not least, GKE. See how it all comes together in this blog. July 2026 Product, technology, and tools updates Product update: Google Cloud Managed Lustre is now GA, and available in four distinct performance tiers that deliver throughput ranging from 125 MB/s, 250 MB/s, 500 MB/s, to 1000 MB/s per TiB of capacity — with the ability to scale up to 8 PB of storage capacity. The Managed Lustre solution is powered by DDN’s EXAScaler, combining DDN's decades of leadership in high-performance storage with Google Cloud's expertise in cloud infrastructure. Product update: C4N network and storage optimized VMs are now GA. C4N is our first network- and block-storage-optimized VM series built to eliminate data-transfer bottlenecks. Powered by 5th Gen Intel Xeon Scalable processors and built on Google's Titanium offloading hardware, it achieves 400 Gbps network bandwidth, 95 million packets per second (MPPS), and up to 25 GiB/s of block storage throughput when paired with Hyperdisk Extreme. New feature: GKE Dataplane V2 up to 15K Nodes with Network Policies (GA). This capability enables standard GKE clusters to scale up to 15,000 nodes while maintaining full active Network Policy enforcement, supporting the massive infrastructure needs of large enterprise and AI/ML customers. New feature: Co-operative time-slicing in llm-d. If you’re running reinforcement learning (RL) workloads, you can now interleave independent RL jobs onto shared physical hardware, increasing aggregate accelerator duty cycles from a ~40% baseline up to 70% without impacting model convergence or accuracy. New AI security tool: Looking to secure your AI supply chain on GKE, deploy AI workloads safely, and cut down on shadow AI? We open-sourced k8s-aibom, a lightweight, unprivileged Kubernetes controller that continuously monitors container clusters to automatically detect running AI runtimes (like vLLM and Triton) and generate standard CycloneDX Machine Learning Bill of Materials (ML-BOMs). Check out the k8s-aibom project and get involved. Practitioner guides and how-tos How-to guide: On July 27, Google announced Day 0 support for Moonshot AI’s Kimi K3 2.8-trillion-parameter open-weight model, the day weights were released. Whichever your preferred deployment path — via Model Garden, custom orchestration, or GKE with llm-d recipes — this guide offers detailed step-by-step instructions to help you evaluate and pilot Kimi K3 in Google Cloud. How-to guide: Google Kubernetes Engine (GKE) managed DRANET supports both GPUs and TPUs. There are several configurations to use this implementation, including standard cluster (where you have full control) and autopilot cluster (where Google does the heavy configs for you). Take a deeper dive in the hands-on lab, GKE Autopilot clusters with TPUs, GKE managed DRANET and Gemma 4. How-to guide: Learn to run Ray on TPUs, not GPUs. In Part 1 of this two-part series, we discuss TPU slices (hint: Ray thinks of them as just another accelerator on which to schedule), then walk through Ray’s various AI libraries (Part 2). How-to guide: Evaluate TPUs for sample workloads using a new microbenchmark suite that helps you accurately assess whether a device is achieving its theoretical performance specifications, and to identify specific performance gaps or architecture-specific bottlenecks. Dive in here. How-to guide: Scale your agents without killing your budget. Learn how GKE orchestration can help you safely pack more agents onto a fixed compute footprint with GKE Agent Sandbox and Pod snapshots. Whether your goal is performance or cost optimization, we teach you how to turn the right dials for optimal agent efficiency. Technical blueprint: Inside the optimization of Mistral 3 large inference on Ironwood. This blog outlines how one Google team optimized Mistral 3 large MoE model inference on Google’s Ironwood (TPU v7x), achieving a 1.5x performance gain. They did so with hybrid sharding, replacing linear VPU summations with tree reductions, optimizing GMM/MLA kernels, and adopting asynchronous scheduling. As a result, they boosted throughput by up to 48% while maintaining benchmark accuracy neutrality. Read the full blog here. Research, reports and deep-dives Report: Google was named a Leader in the inaugural GartnerⓇ Magic Quadrant™ for AI Infrastructure, positioned highest for ‘Ability to Execute’ and furthest for ‘Completeness of Vision’. Gartner called out Google’s proprietary scalable compute, integrated AI Hypercomputer architecture, and the scale of our AI compute capacity as key strengths. Download a copy here. Report: We recently surveyed more than 1,400 senior IT leaders for our State of AI Infrastructure report, and a resounding pattern emerged: The gap between AI ambition and infrastructure reality is widening. In fact, 83% of organizations say they require infrastructure upgrades to support production-grade agentic AI. Read the accompanying blog to understand how adapting your infrastructure to meet the demands that agentic applications place on your systems will help you move from pilot to production. June 2026 Product, technology and tool updates Product update: Protecting sensitive data used with AI is a critical part of advanced and secure cloud infrastructure. Confidential Computing cryptographically protects data in use in hardware-based Trusted Execution Environments (TEEs) with verifiable data integrity, and is now available on the accelerator-optimized G4 machine series, featuring NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. Get started with Confidential G4 VMs and Confidential G4 GKE Nodes. Developer resource: The new TPU Developer Hub is the place to go for model builders, optimizers, and developers to learn to unlock the full performance of Google Cloud TPUs. Read more in this blog. New product: Scale your AI workloads with the new OpenTelemetry-Based TPU AI Telemetry Collector Agent. For the first time, you can route high-fidelity TPU hardware telemetry to Google Cloud Monitoring, Google Managed Prometheus, or your own self-hosted Grafana stack. Practitioner guides and how-tos How-to guide: Learn how to build high availability into an AI inference workload running on GKE Inference Gateway with TPUs, Cloud Storage FUSE and Dynamic Resource Allocation (DRA). This blog provides an overview, or you can get all the technical details in the hands-on codelab. How-to guide: Did you know you can connect your AI agents to unstructured data in Cloud Storage via Model Context Protocol (MCP)? In this blog, learn about why would want to do that from three customer examples, then how to do it, choosing either a fully managed service, or a self-managed local server for more customization and control. Research, reports and deep-dives Report: According to an independent benchmark report, GKE Inference Gateway outperforms the next leading managed Kubernetes service with 15.7% higher throughput, 92.8% shorter wait times, and 62.6% lower inter-token latency. This performance can be attributed to its use of prefix caching, which optimizes LLM performance by storing the KV cache (activation states) of long, repetitive prompt prefixes. Learn more in the blog. Architecture deep dive: A closer look at the cold start problem, this time for TPUs and GKE, and how the Run:ai Model Streamer can help change the dynamic. Customer and partner updates Customer win: Leveraging GKE, BigQuery, Cloud SQL, and Gemini Enterprise Agent Platform, Pager Health is eliminating operational fragmentation to deliver a simplified, personalized U.S. healthcare experience that transforms lives. Customer win: Trustpilot, the customer review platform, built a high-volume streaming pipeline using fine-tuned Gemma models with Dataflow and Gemini Enterprise Agent Platform running on cost-optimized A2 VMs using A100 GPUs, as well as optimized version of vLLM maintained by Gemini Enterprise Agent Platform. May 2026 Product, technology and tool updates Product update: GKE Agent Sandbox is now generally available. New open-source project: Agent Substrate is a new open-source project aimed at continuing to push the limits of agentic infrastructure density New feature: Google AI Edge Portal, a solution for testing and benchmarking on-device machine learning (ML) at scale, now supports benchmarking and debugging on-device LLMs. Read more here. Product deep dive: We went into depth about Cloud Storage Rapid, a new family of high-performance storage offerings for AI workloads. At launch, offerings include Rapid Bucket (formerly Rapid Storage), a high-performance zonal object storage offering, and Rapid Cache (formerly Anywhere Cache), which accelerates reads on-demand and colocates compute and data for workloads in existing buckets. Research, reports and deep dives Architecture deep dive: Google Global Infrastructure VP Bikash Koley and Engineering Fellow Arjun Singh provide a high-level overview of the challenges that AI workloads pose to network infrastructure, and discuss the deep enhancements we’ve made to our data center fabrics, WAN, and global networks to better support them. Architecture deep dive: We unveiled a new cluster-level reliability model for developing frontier AI models on TPUs, ditching instance-level reliability Customer and partner updates Customer win: Visual media provider Imgix serves more than 8 billion images and videos from AI Hypercomputer equipped with G4 VMs powered by NVIDIA RTX PRO 6000 Blackwell GPUs.

Read original article

Google

August 31, 2026

Big Query Graph is now GA: the knowledge foundation for the agentic era

Many of the questions that matter in enterprise data aren't just about individual rows — they're about how things connect: how two accounts are linked, what path a payment took, what context grounds an AI agent's answer. That’s what a graph is built to solve. Historically, unlocking these insights meant extracting data into standalone graph databases, creating silos and operational overhead. To remove these barriers, we brought native graph capabilities directly to the data warehouse. Today, we are announcing the general availability of BigQuery Graph. We introduced BigQuery Graph in preview to unify graph and relational analytics. ISO-standard Graph Query Language (GQL) sits alongside SQL, traversals run natively, and there’s no ETL. And because it’s built on BigQuery, BigQuery Graph inherits and expands its capabilities: It reaches petabyte-scale without the memory bottlenecks of a scale-up database, runs under your existing row- and column-level security, and calls BigQuery ML and AI functions in the same query. One engine, two jobs — large-scale graph analytics, and connected context for AI agents. "BigQuery Graph has been a game-changer for our threat detection pipeline, allowing us to move beyond simple, siloed alerts. By modeling our security signal data as a property graph, we can now perform complex, multi-hop traversals in seconds - something that was previously computationally prohibitive. This graph-centric approach automatically clusters anomalies into coherent attack stories, which, combined with the seamless integration of Gemini models, helps us generate actionable threat narratives. We look forward to integrating native BigQuery Graph algorithms to further streamline our workflows." - Pete Rubio, VP of Global engineering at Thales Cybersecurity Products Since preview, we saw data teams across industries adopt BigQuery Graph for both analytical and agentic workflows: Threat and fraud detection: Security and financial organizations correlate signals across event logs to uncover multi-hop attack paths, fraud networks, and suspicious transaction loops. Supply chain digital twins: Manufacturing and logistics organizations map dependencies across suppliers, parts, and distribution routes to simulate disruptions and optimize fulfillment. Identity resolution and Customer 360: Ad-tech and retail platforms stitch fragmented user identifiers and behavioral touchpoints into unified customer profiles across channels. Knowledge graphs and AI agent grounding: Enterprise AI teams build structured knowledge graphs from unstructured documents, providing domain context to ground Gemini models and GraphRAG workflows. Network lineage and infrastructure management: Telecommunications and enterprise IT teams track complex network topologies, service dependencies, and data lineage across multi-hop paths. What’s new in BigQuery Graph Reaching GA is more than a stability milestone. The work fell into two movements: we made the graph engine itself faster and broader, and we built an agentic ecosystem around it — so agents can build a graph, chat with it, and keep an auditable memory on it. Some of what follows is generally available today; some is in preview or rolling out over the coming weeks. A faster, broader graph engine “Advertising has spent decades optimizing individual events; the agentic era will optimize the relationships between them. At Yahoo, BigQuery Graph gives our AI agents connected context - campaigns, audiences, exposures, and outcomes, traversable with standard GQL right where our monetization data already lives, with no separate graph engine and no data movement. Our agents don't just read the graph; they reason over it and write their conclusions back as new relationships. That's how monetization moves beyond automation, to autonomous systems we can trust to act.” - Mikul Bhatt, Director of Engineering, Monetization Platform at Yahoo Borderless graph Lakehouse Agents are only as good as the context they can reason over, and that context is rarely in one place. With borderless Lakehouse, a single BigQuery Graph can span native BigQuery tables and open Iceberg tables in other clouds — through Databricks Unity Catalog, AWS Glue, or Snowflake — traversed in place, without copying data or building ETL pipelines. Say a support agent needs to answer, "who supplies the product behind this customer's delayed order, and where are they based?" The customer data sits in an Iceberg lakehouse on Google Cloud, the product and supplier records in a Databricks catalog on AWS. Instead of stitching the sources together per request, the agent traverses one virtual knowledge graph that already connects them — over data that never moved. Figure 1: A diagram illustrating a virtual knowledge graph spanning across Google Cloud (blue nodes), AWS (yellow nodes), and other clouds (green nodes) without data movement. The following DDL statement shows how you can define this virtual graph, mapping your node and edge tables directly across both cloud environments: code_block <ListValue: [StructValue([('code', '-- A virtual knowledge graph spanning two clouds - no data movement\r\nCREATE OR REPLACE PROPERTY GRAPH `my_project.retail.virtual_kg`\r\n NODE TABLES (\r\n -- Google Cloud\r\n `my_project.gcs_lake.retail.customers` AS Customer KEY (customer_id),\r\n -- AWS\r\n `my_project.dbx_fed_catalog.retail.products` AS Product KEY (product_id),\r\n `my_project.dbx_fed_catalog.retail.suppliers` AS Supplier KEY (supplier_id)\r\n )\r\n EDGE TABLES (\r\n `my_project.gcs_lake.retail.purchases` AS Bought KEY (purchase_id)\r\n SOURCE KEY (customer_id) REFERENCES Customer (customer_id)\r\n DESTINATION KEY (product_id) REFERENCES Product (product_id),\r\n `my_project.dbx_fed_catalog.retail.products` AS Supplied_By KEY (product_id)\r\n SOURCE KEY (product_id) REFERENCES Product (product_id)\r\n DESTINATION KEY (supplier_id) REFERENCES Supplier (supplier_id)\r\n );'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f7db8f575b0>)])]> With that, the agent gets a grounded, multi-hop answer assembled across two clouds in a single traversal: code_block <ListValue: [StructValue([('code', "-- Agent grounding: trace a customer to the supplier behind their product, across clouds\r\nGRAPH `my_project.retail.virtual_kg`\r\nMATCH (c:Customer {customer_id: 'C1'})-[:Bought]->\r\n (:Product)-[:Supplied_By]->(s:Supplier)\r\nRETURN s.name AS supplier, s.country AS supplier_country"), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7f7db86f5af0>)])]> Faster and more expressive GQL BigQuery Graph is built for questions about connection: how two accounts are linked, what path a payment took, which entities sit within a few hops of a flagged one. These are the questions SQL joins struggle to express, and they're where a graph engine earns its place. At GA, we've made them both faster to run and easier to write: Faster execution. GA optimizes path-finding for acyclic and undirected traversals: against public benchmarks, GQL is 2x faster since preview and undirected traversal 100x, with faster, more resource-efficient cycle detection in ACYCLIC and TRAIL path modes. Lower query latency keeps the neighborhood and path lookups that ground an agent's answer responsive under frequent, interactive access. More expressive queries. With the new CALL statement and extended subquery support, you can run a graph subquery for each entity in a result, or invoke a reusable named function, so a complex question breaks into parts instead of one sprawling pattern. The same functions an analyst writes become the building blocks an agent calls as a tool. Built for the agentic era “Companies have plenty of workforce data, but very little shared understanding of what their people can do or where they fit. BigQuery Graph lets us turn that scattered information into a reusable property graph and traverse connections across people, roles, capabilities, and evidence at scale, so the same connected workforce context can support thousands of decisions instead of being recreated one decision at a time. That gives AI a stronger foundation for much harder questions about how work should get done.” - Heiko Roth, Founder & CEO, Workerbee Chat with your graphs You don't have to write GQL to explore a graph. BigQuery conversational analytics lets you chat with your graph directly in natural language: it reads the relationships in your schema to translate a question into SQL or GQL, and visualizes the traversal for path-based answers. The agent draws on graph metadata like descriptions and synonyms to keep results grounded — the relationships that make a graph a graph are exactly what cut the ambiguity and hallucination that plague free-form natural language querying. You can also connect Gemini Enterprise to BigQuery Graph through an MCP server, or publish the conversational data agent to it directly. Build a graph with an agent Standing up a graph — modeling tables into nodes and edges, then writing GQL against them — is work you can hand to the data agent you already use. We've packaged BigQuery Graph expertise into an agent skill that makes your agent fluent in graph: GQL pattern matching, blending graph and SQL, and schema design that follows our recommended practices. The capabilities are accessible out of the box in your preferred agentic coding tool, such as Antigravity, Visual Studio Code, Claude Code, and Codex, with the Google Cloud Data Agaent Kit extension. The skill is also learning to author, not just advise — a capability rolling out soon. Point it at a dataset, a model document, or an ER diagram and it proposes the nodes and edges, then verifies each relationship against your data before building, showing you the match rates: this one resolves at, say, 98%, that one 56%. You get a graph you can trust from day one. Give your agents an auditable memory Grounding an agent is half the job; the other half is remembering what it did. As agents move from advising to acting, every decision has to be explainable after the fact — which option was chosen, which policy applied, which alternatives were rejected. With context graph in BigQuery Agent Analytics, each action an agent takes is captured and shaped into a context graph: a typed, queryable trace of the agent's reasoning, stored right in BigQuery Graph. Because the trace is itself a graph, "why did the agent do this?" is a single traversal — and the outcomes you join back to those decisions become the data that improves the next one. Get started with BigQuery Graph today BigQuery Graph runs graph analytics and grounds AI agents on your data, across clouds. To get started, check out the overview and data model to see how GQL, node tables, and edge tables fit together, then put them to work on your team’s common patterns. Trace suspicious money movement and synthetic identities in the fraud detection codelab, stitch fragmented emails, devices, and cookies into one customer in the identity resolution codelab, or model a supply chain as a digital twin you can query for hidden dependencies when disruption hits. From there, take it toward agents. The agent context graph codelab turns raw event logs into a graph that audits, explains, and traces what your autonomous agents actually did — the connected memory behind a system you can trust to act. If your workloads span both real-time operational transactions and massive-scale analytics, explore our unified graph solution to see how Spanner Graph and BigQuery Graph work together. And when you are ready to go deeper — our ebook walks the journey end-to-end.

Read original article

Nvidia

August 31, 2026

On the CUBE Pod: Nvidia steamrolls expectations as Mythos shakes up cybersecurity

Nvidia Corp. might as well be swimming in money like Scrooge McDuck. The artificial intelligence firm had yet another astounding quarter, beating expectations for revenue. With that announcement came a slew of reports: Nvidia’s expanded collaboration with Cisco Systems Inc. on AI infrastructure, its reported $12.9 billion deal to acquire Hugging Face and another investment […] The post On theCUBE Pod: Nvidia steamrolls expectations as Mythos shakes up cybersecurity appeared first on SiliconANGLE.

Read original article

OpenAI

August 31, 2026

Open AI says its Chat GPT ad business hits a $1 billion annual run rate

OpenAI's advertising business has reached an annualized revenue run rate of $1 billion, according to the company. The article OpenAI says its ChatGPT ad business hits a $1 billion annual run rate appeared first on The Decoder.

Read original article

Nvidia

August 31, 2026

Nvidia’s $3.5 B Media Tek bet reveals its plan for tackling Big Tech’s AI chip buildout

Nvidia invests $3.5 billion into Taiwanese chipmaker MediaTek. The deal shows how Nvidia plans to stay essential to AI infrastructure as Big Tech begins to build its own AI chips.

Read original article

Nvidia

August 31, 2026

We Accelerate Every AI Model in the World, Huang Says

Nvidia CEO Jensen Huang says their chips accelerate the entire lifecycle of artificial intelligence. He speaks exclusively with Bloomberg's Ed Ludlow and says Nvidia can address markets that no one else can. (Source: Bloomberg)

Read original article

OpenAI

August 31, 2026

Chat GPT now faces stricter EU oversight as a very large search engine

The EU Commission is classifying ChatGPT as a very large search engine under the Digital Services Act for the first time, with at least 45 million monthly EU users. By the end of 2026, OpenAI has to deliver risk assessments, transparency reports, and an ad archive, among other things. Whether the Commission can also demand access to training data is disputed among legal experts. The article ChatGPT now faces stricter EU oversight as a very large search engine appeared first on The Decoder.

Read original article

OpenAI

August 31, 2026

Open AI starts charging some customers only when its AI actually works

OpenAI is offering some large customers outcome-based pricing, where they pay only once the AI actually finishes a task. Salesforce, Adobe, and several startups are also moving away from fixed subscription fees. The central dispute stays the same. Who gets credit for the success, the software or the customer? The article OpenAI starts charging some customers only when its AI actually works appeared first on The Decoder.

Read original article

Nvidia

August 31, 2026

Nvidia Makes $3.5 Billion Bet on Media Tek

Nvidia is investing $3.5 billion in MediaTek, deepening collaboration with the Taiwanese chipmaker as it's working to persuade more companies to build chips that plug into its dominant data center ecosystem. Nvidia CEO Jensen Huang and MediaTek CEO Rick Tsai join Bloomberg's Ed Ludlow in this exclusive interview to discuss the deal. (Source: Bloomberg)

Read original article

Nvidia

August 31, 2026

Nvidia Makes Media Tek Partnership Even Bigger, Huang Says

Nvidia Corp. Chief Executive Officer Jensen Huang and Rick Tsai, MediaTek Inc. CEO, talk about expanding their partnership. Nvidia is investing $3.5 billion in the Taiwanese chipmaker. They speak exclusively to Bloomberg's Ed Ludlow. (Source: Bloomberg)

Read original article

Microsoft

August 31, 2026

Fireflies AI co-founder and CEO Krish Ramineni shares the ‘biggest problem’ with Google, Microsoft meetin - The Times of India

Fireflies AI co-founder and CEO Krish Ramineni shares the ‘biggest problem’ with Google, Microsoft meetin The Times of India

Read original article

Anthropic

August 31, 2026

Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time.

Anthropic is putting AI agents to work on one of the field’s hardest problems: keeping other AI systems aligned with The post Anthropic’s Claude fixed all 10 alignment failures. Then it tried to cheat 2.4% of the time. appeared first on The New Stack.

Read original article

Nvidia

August 31, 2026

Nvidia Investing $3.5 Billion in Chipmaker Media Tek

Nvidia says it's investing $3.5 billion in Taiwanese chipmaker MediaTek. Nvidia will buy bonds that convert into MediaTek shares. Bloomberg's Ed Ludlow reports. (Source: Bloomberg)

Read original article

OpenAI

August 31, 2026

Open AI-led coalition warns AI will compress cyberattack timelines, expose enterprise weaknesses - csoonline.com

OpenAI-led coalition warns AI will compress cyberattack timelines, expose enterprise weaknesses csoonline.com

Read original article

Anthropic

August 31, 2026

Anthropic Legal Fight With Pentagon Shows Shifting Politics of AI

Judge says Pentagon move was made in retaliation for company’s views

Read original article

Meta

August 31, 2026

Pocket's AI made my game ideas real. Now Meta controls the results.

Interactive mobile "gizmos" are easy to make, hard to share outside Meta's platform.

Read original article

Google

August 31, 2026

Hollywood studios have sued over AI. Now Google is courting them to use its tools - Los Angeles Times

Hollywood studios have sued over AI. Now Google is courting them to use its tools Los Angeles Times

Read original article

Anthropic

August 31, 2026

China Rebukes Anthropic, Sets Terms for Key US-China AI Dialogue - Bloomberg.com

China Rebukes Anthropic, Sets Terms for Key US-China AI Dialogue Bloomberg.com

Read original article

Anthropic

August 31, 2026

Open AI and rival AI labs are buying tens of thousands of Mac minis to train computer-use agents

According to The Information, OpenAI has purchased tens of thousands of Mac minis and Mac Studios to train computer agents. Anthropic also relies on Apple hardware. Demand is so high that the most powerful models have been sold out for months. Apple’s Mac revenue rose by nearly 29 percent to $10.4 billion in the June quarter. The article OpenAI and rival AI labs are buying tens of thousands of Mac minis to train computer-use agents appeared first on The Decoder.

Read original article

Hugging Face

August 31, 2026

Normas TCU --- A Brazilian Portuguese IR Dataset and an Evaluation of LLM-as-a-Judge for Relevance Assessment

arXiv:2608.27746v1 Announce Type: new Abstract: Portuguese Information Retrieval (IR) lacks public datasets, and relevance assessment for specialized collections remains costly. While Large Language Models (LLMs) increasingly support relevance assessment, their reliability in non-English specialized domains remains unclear. We introduce NormasTCU (https://huggingface.co/datasets/LeandroRibeiro/NormasTCU), a Brazilian Portuguese IR dataset with 14,469 legal documents, 46 queries, and 3,048 human judgments over 812 query-document pairs. Using NormasTCU, we evaluated LLM-as-a-judge for relevance assessment by prompting three models with two prompt techniques to grade these pairs. We then compared the rankings of 15 IR systems derived from LLM-generated and human reference qrels. LLMs consistently showed a positive scoring bias (mean absolute error: 0.46--0.66 on a 0-2 scale). Furthermore, pair-level agreement with human judgments achieved only fair to moderate levels, with Cohen's kappa ranging from 0.32 to 0.53. Despite this bias, LLM-generated judgments often yielded highly similar system rankings for nDCG@10 and MRR (observed Kendall's tau greater than or equal 0.90, although the bootstrap confidence intervals did not always remain above this threshold), but were less reliable for P@10 and R@10. Notably, LLM-based rankings were sometimes more strongly correlated with the reference ranking than individual human annotations were. As a practical implication, our results suggest that LLMs could effectively support scalable relevance assessment in specialized Portuguese corpora when evaluated using nDCG or MRR (rank-aware metrics), but they should be avoided when relying on precision or recall.

Read original article

Nvidia

August 31, 2026

GPU-Native Approximate Nearest Neighbor Search with IVF-Ra Bit Q: Fast Index Build and Search

arXiv:2602.23999v2 Announce Type: replace Abstract: Approximate nearest neighbor search (ANNS) on GPUs is gaining increasing popularity for modern retrieval and recommendation workloads that operate over massive high-dimensional vectors. Graph-based indexes deliver high recall and throughput but incur heavy build-time and storage costs. In contrast, cluster-based methods build and scale efficiently yet often need many probes for high recall, straining memory bandwidth and compute. Aiming to simultaneously achieve fast index build, high-throughput search, high recall, and low storage requirement for GPUs, we present IVF-RaBitQ (GPU), a GPU-native ANNS solution that integrates the cluster-based method IVF with RaBitQ quantization into an efficient GPU index build/search pipeline. Specifically, for index build, we develop a scalable GPU-native RaBitQ quantization method that enables fast and accurate low-bit encoding at scale. For search, we develop GPU-native distance computation schemes for RaBitQ codes and a fused search kernel to achieve high throughput with high recall. With IVF-RaBitQ implemented and integrated into the NVIDIA cuVS Library, experiments on cuVS Bench across multiple datasets show that IVF-RaBitQ offers a strong performance frontier in recall, throughput, index build time, and storage footprint. For Recall approximately equal 0.95, IVF-RaBitQ achieves 3.0x higher QPS than the state-of-the-art graph-based method CAGRA, while also constructing indices 14.7x faster on average. Compared to the cluster-based method IVF-PQ, IVF-RaBitQ delivers on average over 4.5x higher throughput while avoiding accessing the raw vectors for reranking.

Read original article

Meta

August 31, 2026

Safe Step: An Interactive Demonstration of Semantic Communication for Pedestrian Safety Monitoring

arXiv:2608.27688v1 Announce Type: new Abstract: In this paper, we develop SafeStep, an interactive browser-based semantic communication platform for live pedestrian safety monitoring. SafeStep extracts pedestrian information from four live traffic-camera feeds, transmits it through a semantic communication transceiver over an Additive White Gaussian Noise (AWGN) channel, and renders user-specific positions, trajectories, and risk labels. The platform allows to independently select the transceiver, Signal-to-Noise Ratio (SNR), codelength, and Age of Information (AoI), and demonstrates the transceiver performance of the selected configuration through live pedestrian safety monitoring to each browser. SafeStep compares a recently proposed semantic communication design called Meta-VIB with five baseline transceivers. Meta-VIB uses a compact neural model with only $4.16$ million parameters to generalize across varying SNR, codelength, and AoI values without online retraining. Experimental results show that Meta-VIB achieves mean task-loss reductions of up to $92.1\%$. On one high-end GPU server, the integrated concurrent-access workload maintains the target $5$ frames/s through $20$ users. At $100$ users, each requesting a distinct configuration, SafeStep records no request failures and a mean application response time below $1$ s, but its mean per-browser frame rate falls to approximately $1$ frame/s. To our knowledge, SafeStep is the first real-time semantic communication platform to make AoI-induced downstream degradation directly observable in live monitoring applications.

Read original article

OpenAI

August 31, 2026

A milestone in expanding access to AI

ChatGPT Ads reaches $1 billion in annualized revenue run rate and expands globally, supporting broader access to AI through free and affordable options.

Read original article

OpenAI

August 31, 2026

Exclusive | The $5.5 Billion Perk Soft Bank’s Data-Center Venture Offered to Land Open AI - wsj.com

Exclusive | The $5.5 Billion Perk SoftBank’s Data-Center Venture Offered to Land OpenAI wsj.com

Read original article

Anthropic

August 31, 2026

Claudeforce Could Put Anthropic's AI in Front of Millions of Salesforce Users - Memeburn

Claudeforce Could Put Anthropic's AI in Front of Millions of Salesforce Users Memeburn

Read original article