Daily News
AI research, safety, product, and engineering links
Latest AI Reading
Source-dated posts from the last 14 days
Updated 2026-08-07 23:35 UTC
TutorMoments: Do AI tutors know when to help and when to hold back?
Read originalHow Cohere Health digitizes clinical policies using Amazon Bedrock AgentCore
In this post, you learn how Cohere Health built a multi-tenant agentic architecture on AgentCore using AgentCore Runtime’s secure MicroVM isolation, unified tool access through AgentCore Gateway, AgentCore Memory, and t...
Read originalHow TReNDS automates root-cause analysis with Amazon Bedrock
TReNDS, a research center at Georgia State University, built an agentic AI pipeline on Amazon Bedrock and the open-source Strands Agents SDK that automatically investigates production errors in real time, reducing root-...
Read originalDetermining playoff clinching scenarios in the NHL using constraint programming
The AWS Generative AI Innovation Center built an automated system that uses constraint programming and custom tree search to determine, with mathematical certainty, when and how an NHL team clinches a playoff spot. The...
Read original[AINews] AMD buys Taalas
The Inference Inflection is HEATING up.
Read originalScaling Categorical Flow Maps
Continuous diffusion and flow matching models could represent a powerful alternative to autoregressive approaches for language modelling (LM), as they unlock a host of advantages currently reserved for continuous modali...
Read originalBeyond Next-Token Prediction: A Performance Characterization of Diffusion versus Autoregressive Language Models
Large Language Models (LLMs) have achieved state-of-the-art performance on a broad range of Natural Language Processing (NLP) tasks, including document processing and code generation. Autoregressive Language Models (ARM...
Read originalArbitrage: Efficient Reasoning via Advantage-Aware Speculation
Modern Large Language Models achieve impressive reasoning capabilities with long Chain of Thoughts, but they incur substantial computational cost during inference, and this motivates techniques to improve the performanc...
Read originalWhy do models task game?
TL;DR How can we study misalignment with today's models as proxies? They're clearly not paperclip maximizers, but they also often do things the user doesn't want. A strong contender for a real misaligned propensity is t...
Read originalNVIDIA Vera Storage Benchmarks: Faster Encryption, Compression, Integrity Checking, and Recovery for AI-Native Storage
Storage is an active part of every agentic AI workflow. As agents retrieve enterprise knowledge, access persistent memory, reuse key-value (KV) cache data,...
Read originalNVIDIA Nemotron 3 Ultra Leads Open Models on Accuracy and Efficiency in Agentic RTL Coding
Modern chip design is increasingly limited by engineering time. Register transfer level (RTL) development and verification require specialized hardware...
Read originalAdvancing Semiconductor Innovation Across Materials Engineering and Manufacturing
As AI workloads increase, explosive compute demand is pushing the semiconductor industry to meet unprecedented performance targets. Even small delays can have...
Read originalSix Agent Harness Capabilities for Higher Model Performance
Building a great AI agent isn’t just about choosing the right models. The harness is the architecture surrounding the model. How it renders context, executes...
Read originalNVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they...
Read originalDeveloping Healthcare Robotics with GPU-Native Medical Physics Simulation
Unlike autonomous driving or industrial robotics, healthcare robotics can’t rely on internet-scale data collection or unlimited real-world experimentation....
Read originalHow to Self-Host a Validated AI Coding Assistant with NVIDIA NeMo Guardrails
Deploying an AI coding assistant in a regulated, sovereign, or source-sensitive environment, often comes with challenges. Three common issues are: the source...
Read originalNVIDIA Exemplar Cloud: Lessons for Unlocking Full Performance on AI Infrastructure
Two AI computing clusters built from identical NVIDIA H100, GB200 NVL72, or GB300 NVL72 systems can deliver materially different training throughput. We...
Read originalFour Ways to Deploy More Secure AI Agents
Knowledge workers are increasingly integrating AI agents into their workflows. Agents that function as "digital coworkers" offer clear benefits. For example,...
Read originalRun High-Performance Core Math at Scale with NVIDIA nvmath-python
NVIDIA nvmath-python is a library designed to bridge the gap between the Python scientific community and NVIDIA CUDA-X math libraries. It gives Python users...
Read originalSecuring AI agents with temporal policies in Amazon Bedrock AgentCore
Temporal policies in Amazon Bedrock AgentCore let you define stateful rules that evaluate authorization based on an agent's session history. Learn how to enforce workflow sequencing, prevent data fabrication, cap financ...
Read originalConfigure rate limits for AI traffic on AgentCore gateway
Learn how to configure rate limits on Amazon Bedrock AgentCore gateway to enforce per-user and per-target traffic controls. Define request, token, and connection limits scoped by JWT claims or IAM identity to protect do...
Read originalControl agent behaviors and cost beyond a single action: new capabilities in Amazon Bedrock AgentCore
Learn about new capabilities in Amazon Bedrock AgentCore: temporal policies powered by Dogwood, a new open source policy language for AI agents, and rate limiting on the gateway. These features give you deterministic co...
Read originalBuild visibility for Codex on Amazon Bedrock with OpenTelemetry and Amazon CloudWatch
As engineering teams adopt coding agents like Codex, leaders need visibility into adoption, consumption, and reliability. This post shows how to route Codex OpenTelemetry metrics through a local collector to Amazon Clou...
Read originalEnforcing data residency with single-Region Claude Code on Amazon Bedrock
A regulated customer needed all Claude Code inference processed in a single AWS Region (London), not just in-geography. This post shows two ways to pin Claude Code on Amazon Bedrock to one Region: an application inferen...
Read originalAgent Skills for Automated Reasoning policies in Amazon Bedrock
Learn how to run the full Amazon Bedrock Automated Reasoning policy lifecycle from your coding agent. A suite of open source Agent Skills builds, reviews, tests, debugs, deploys, and validates a custom policy end to end...
Read originalBuilding an agentic app deployer with Amazon Bedrock and AWS Lambda
PDI Technologies built PDI Brew, an agentic platform on AWS where non-technical employees describe a tool in plain English and receive a fully provisioned, multi-tenant web application in seconds. See how a pluggable pl...
Read originalWeatherNext: AI model achieves breakthrough in forecasting cyclones
Read originalGeForce NOW Shakes Up August With 26 New Games
August is here, bringing 26 new games for GeForce NOW members. Command the seas in World of Warships: Legends and discover what’s next in the GeForce NOW library, starting with the eight newly added games this week. In...
Read originalInto the Omniverse: How Open World Models Push the Frontier of Physical AI
In July, NVIDIA joined more than 200 companies and organizations in signing “Open Weights and American AI Leadership,” an open letter arguing that AI leadership will be measured not by any single frontier model but by w...
Read original[AINews] Jeff, Sanjay, Oriol, and Quoc depart DeepMind; Demis to Chair; Koray to SVP — what is going on at GDM???
The end of an era.
Read originalBaseten on Hugging Face Inference Providers 🔥
Read originalLocking Pretrained Weights via Deep Low-Rank Residual Distillation
The quality of open-weight language models has dramatically improved in recent years. Sharing weights greatly facilitates model adoption by enabling their use across diverse hardware and software platforms. They also al...
Read originalDeepAmbigQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness
Large language models (LLMs) with integrated search tools show strong promise in open-domain question answering (QA), yet they often struggle to produce complete answer set to complex questions such as âWhich actor fr...
Read originalNVIDIA and Partners Build in America, for America
NVIDIA and its partners are investing in American manufacturing, supply chains, energy grids and skilled workforces so the U.S. can produce the infrastructure needed for better healthcare, breakthrough scientific discov...
Read original[AINews] Megakernels are so dead and so back
A quiet day lets us highlight a Cursor launch and an engineering debate
Read originalTaming Outlier Tokens in Diffusion Transformers
We study outlier tokens in Diffusion Transformers (DiTs) for image generation. Prior work has shown that Vision Transformers (ViTs) can produce a small number of high-norm tokens that attract disproportionate attention...
Read originalReturning to ARC
I've returned to the Alignment Research Center (ARC) as executive director. My main focus for the next six months will be driving forward ARC's research agenda—building techniques to find mechanistic explanations for ne...
Read originalUnpacking ChatGPT Work: the Agent for a Billion Users
An external reconstruction of how Memory, Proactivity, Scheduling, Browser Use, Plugins, Skills and Tools work in the new ChatGPT Work.
Read originalNVIDIA Joins NSF State and Regional AI Hubs Program to Expand AI Research and Education Across the US
NVIDIA is participating in the U.S. National Science Foundation’s (NSF) State and Regional Artificial Intelligence Infrastructure Hubs program, an effort launching today to expand access to the advanced computing, data,...
Read originalNVIDIA Alpamayo 2 Super, the Frontier Open Model for Robotaxis and Autonomous Vehicles, Now Available for Commercial Use
For robotaxis and other autonomous vehicles (AVs), the hardest problems aren’t the everyday scenarios. They’re the rare, complex situations that are difficult to anticipate and train for. Handling these long‑tail events...
Read originalAs AI Increases Demands on Memory, Storage Steps Up
Surging AI demands are driving the need for massive datasets and context windows that burst past the confines of system memory. But rising needs aren’t met by simply adding more storage capacity. What’s needed is useful...
Read originalDeploy local agents everywhere with LFM2.5-2.6B
Read originalAI Leaders Propose SAFE Guidelines for Cybersecurity Transparency
Members of the Open Secure AI Alliance — now more than 120 organizations strong — are developing new guidelines to strengthen agentic AI cybersecurity as the annual Black Hat conference begins in Las Vegas today. The Li...
Read originalIntroducing Shieldstral.
Shieldstral introduces a 3B open-weights multimodal safety classifier that outperforms models up to 7x its size.
Read original[AINews] Qwen 3.8 Max(2.4T) and 27B, new open weights models for Coding and Cowork
Qwen is so back!
Read originalThe Inference Engineering Masterclass — Philip Kiely & Ali Taha, Baseten
Baseten just raised a $13B Series F and is now one of the leading kings of inference engineering. We go into everything you need to know for autoregressive and diffusion engineering.
Read originalOrchard: An open framework for scalable agentic AI
Orchard is an open-source framework for the research community to train and evaluate AI agents across task types. It reduces complexity while supporting strong performance from smaller models by enabling researchers to...
Read originalIntroducing our Artifacts Hub and Adoption Dashboard
Scaling our curation and measurement of the open ecosystem.
Read originalImport AI 467: Self-sustaining AI viruses; pacing AI progress; confusion about AI and creativity
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Self-sustaining and self-replicating...
Read originalConcrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face
Three-Minute Executive Summary An OpenAI model/multi-agent system bypassed its sandbox and launched a cyberattack on Hugging Face in order to cheat on a cyber evaluation. In this post, we describe the ambitious, compreh...
Read originalHow we built a realtime system for responsive voice AI in six months
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
Read originalUnderstanding Alignment in Multimodal LLMs: A Comprehensive Study
Preference alignment has become a crucial component in enhancing the performance of Large Language Models (LLMs), yet its impact in Multimodal Large Language Models (MLLMs) remains comparatively underexplored. Similar t...
Read originalLatest open artifacts (#23): Laguna S2.1, Inkling, & Kimi K3 show the utility of open models on the Pareto frontier
Capacity to train strong models is proliferating.
Read original[AINews] not much happened today
apart from DeepSeek V4-Flash 0731, a quiet day.
Read originalSOTA alignment assessments don’t strongly update us against misalignment
Anthropic concluded in the April Mythos Preview alignment risk update that the model “does not possess any unknown propensities that would increase alignment risk.” The report argues that if Mythos Preview were coherent...
Read originalValue Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values
TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is t...
Read originalAGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)
Cross-posted from our new Substack It’s been nearly two years since our last major update here in August 2024 and we wanted to share another recap of our recent work with the AGI safety community. Things have changed a...
Read originalThe AGI Safety and Alignment team at Google DeepMind is Hiring (July 2026)
GDM’s AGI Safety and Alignment Team is hiring for multiple roles, across all areas in this post on our recent work . This is the team at GDM, led by Rohin Shah , that aims to reduce existential risks from AI systems. Yo...
Read originalOpenAI has already ended an internal pause
One day before OpenAI’s HF incident disclosure, OpenAI disclosed that it paused internal deployment of a long-horizon model after it circumvented its sandbox, then restored access weeks later under new monitoring. So a...
Read original[AINews] GPT 5.6 price cut by 20%-80%: Cost of GPT 5.4 Intelligence dropped 13x in 4 months due to GPT 5.6 recursive self-optimization
Distillation is all you need!
Read originalPromising Signals on AI Governance from China
View the official memo here. China has consistently signaled a willingness to engage on global AI governance since at least 2017. This memo compiles key statements from the Chinese government and prominent figures demon...
Read originalScience One Framework: A verifiable autonomous research framework via Chain-of-Evidence
General Science
Read originalEchoverse: Deep, evolving environments for computer-use agents
Computer-use AI agents struggle with multi-step workflows like email and customer support. Echoverse trains agents in realistic environments rather than simply providing more training tasks, helping them improve as the...
Read originalEvoLib: Turning experience into evolving knowledge
LLMs do not get smarter just by remembering more. EvoLib turns experience into evolving knowledge, taking reusable skills and insights that help models learn and adapt across tasks long after deployment. The post EvoLib...
Read originalGPU Management: Why Idle GPUs Are the New Grounded Aircraft
Read originalGemini Robotics ER 2: powering robotics with video understanding, task orchestration, and multi-robot collaboration
Gemini Robotics ER 2 helps robots reason, collaborate, and solve real-world tasks. It represents a step change in video understanding, tool orchestration, and multi-robot collaboration for robotic applications.
Read originalThousand-dimensional structure
Summary: One area we plan to explore at Resolution is personas and character training, operationalized as finding and controlling low-dimensional structure in models that emerges in pretraining and flows through post-tr...
Read originalBest in Class: Stream PC Games and Study on the Same Laptop With GeForce NOW
Back to school means balancing assignments, deadlines and downtime. GeForce NOW makes it easy to have it all. With cloud gaming, everyday laptops used for class can also become GeForce RTX-powered gaming setups. When it...
Read originalOntologies Are So Back: Why AI Agents Are Reviving the Semantic Web
AI engineers are rediscovering ontologies as a way to keep probabilistic agents inside deterministic boundaries.
Read originalMoMo: Dial Motion Mode in Robot Manipulation with Spatiotemporal Action Tokenization
To operate effectively across diverse contexts, robots must not only perform manipulation tasks accurately but also adapt how their actions unfold to the task, object, and interaction setting. We ask whether this execut...
Read originalDimensionality Reduction Meets Network Science: Sensemaking on UMAPâs kNN Graph
While UMAP is widely used for exploring high-dimensional data, typical workflows focus on its lower-dimensional embedding, largely overlooking the rich k-nearest-neighbor (kNN) graph that UMAP constructs internally. Thi...
Read original[AINews] AI is eating Finance; AIE NYC now open
a quiet day lets us cover how AI is permeating financial services as the next big vertical after coding.
Read originalImprecise beliefs: a tiny introduction
Richard Ngo challenged me to set a time box and write down as many of the most important features of my formal epistemology as I can in one sitting. Here goes. Motivation: Where probability distributions fail... ...to e...
Read originalWeâre launching Lyria 3.5 in Google Flow Music, with advances across musicality, lyrics, vocals, and creative control
Read originalValue Generalisation 3: Pre-aligned AIs
When we get explicit strong generalisation to work (see the first post on the matter and the second ) my dream would be to create pre-aligned generalising AIs. Think about the usual conflict between alignment and capabi...
Read originalFrom CUDA to MLX: How K-Search Brings Decades of Kernel Expertise to Apple Silicon
Figure 1: CUDA-to-MLX optimization translation map. CUDA optimization knowledge can be translated into architecture-native MLX strategies rather than copied instruction-for-instruction. We face a new epoch in computing....
Read originalHow GPT-5.6 fuses frontier intelligence with frontier efficiency
GPT-5.6 improves AI efficiency across models, inference, and agentic workflows, helping deliver more useful intelligence per dollar.
Read originalPowerful Compute So Compact, It’s Clutch — Build AI Anywhere With NVIDIA Jetson
As a discerning AI investor who values style and substance, Sarah Guo knows this season’s standout accessory isn’t the latest designer purse — but what’s inside it. In a recent video, Guo, founder of AI-native venture c...
Read originalGemini Robotics 2 brings whole body intelligence to robots
Read originalHow independent researchers could investigate AI propensities after misalignment incidents
AI agents sometimes autonomously take sophisticated, sustained actions in clear violation of user and developer intent. As an example, last week OpenAI reported that some of its internal frontier agents autonomously hac...
Read originalSources
Frontier Labs
OpenAI Alignment Research Blog
OpenAI alignment research.
OpenAI Engineering
OpenAI engineering articles and system-building notes.
Anthropic Research
Anthropic research on alignment, interpretability, evaluations, and societal impacts.
Anthropic Engineering
Anthropic engineering team posts and system-building articles.
Anthropic Alignment Science
Anthropic alignment science, interpretability, and risk evaluation articles.
Google DeepMind Blog
Google DeepMind official blog.
Google Research Blog
Official Google Research blog.
Meta AI Blog
Meta AI official blog.
Microsoft Research Blog
Microsoft Research blog.
AI Safety and Governance
LessWrong
AI alignment, rationality, and AI risk discussion.
AI Alignment Forum
Technical AI alignment research community.
AI Alignment
Alignment essays and research posts.
MIRI Blog
Machine Intelligence Research Institute blog.
METR
Model evaluation and frontier-risk research.
METR Evaluations
METR evaluation reports.
Apollo Research Blog
Frontier AI risk, scheming, and evaluations.
Redwood Research Blog
AI risk and safety research.
FAR.AI Blog
AI safety and alignment research.
CAIS Blog
Center for AI Safety updates.
Goodfire Blog
Interpretability and model control.
Goodfire Research
Goodfire research index.
Epoch AI Blog
AI trends, compute, data, economics, and forecasting.
Epoch AI Latest
Unified stream for papers, newsletters, data insights, and podcasts.
AI Companies and Research Labs
Hugging Face Blog
Open-source models, Transformers, applications, and research.
Hugging Face Daily Papers
Daily AI paper discovery.
NVIDIA Technical Blog
GPU, CUDA, AI systems, and technical engineering posts.
NVIDIA Blog
NVIDIA news and applications.
Amazon Science Blog
Amazon research posts, including AI and machine learning.
AWS Machine Learning Blog
AWS ML engineering, product, and practice posts.
Apple Machine Learning Research
Apple machine learning research.
Cohere Blog
Cohere official blog.
Cohere Research
Cohere research posts.
Mistral AI News
Mistral official news and releases.
xAI News
xAI official news.
Academic Labs
BAIR Blog
Berkeley AI Research blog.
Stanford AI Lab Blog
Stanford AI Lab blog.
Stanford HAI News
Stanford HAI news and blog posts.
MIT CSAIL News
MIT CSAIL news.
MIT LINGO Blog
MIT Language and Intelligence group blog.
Personal Blogs and Newsletters
Lilian Weng - Lil'Log
Long-form posts on reinforcement learning, LLMs, agents, and alignment.
Jay Alammar Blog
Visual explanations for machine learning and Transformers.
Andrej Karpathy - Bear Blog
Andrej Karpathy's newer blog.
Andrej Karpathy - Old Blog
Older Karpathy blog posts.
Sebastian Ruder Blog
NLP and machine learning blog.
Sebastian Ruder Newsletter
NLP news and newsletter.
Import AI
Jack Clark's AI research and industry newsletter.
Import AI Substack
Substack version of Import AI.
Interconnects
Nathan Lambert's frontier AI research and industry newsletter.
Latent Space
AI Engineer newsletter and podcast.
DeepLearning.AI - The Batch
AI news digest.
Distill
Classic visual and explanatory machine learning articles.
The Gradient
AI research, society, and commentary.
