The Archive
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Accelerating Long-Context Model Training in JAX and XLA
Large language models (LLMs) are rapidly expanding their context windows, with recent models supporting sequences of 128K tokens, 256K tokens, and beyond.... Large language models (LLMs) are rapidly expanding their context windows, with recent models supporting sequences of 128K tokens, 256K tokens, and beyond. However, training these models with extended context lengths presents significant computational and communication challenges. As context lengths grow, the memory and communication overhead of attention mechanisms scale quadratically… Source
Defining AI automation: A new kind of workplace
Enterprise AI strategy template and guidance on choosing automation approaches; generic business advisory content.
The Sora feed philosophy
OpenAI outlines Sora feed design philosophy emphasizing personalized recommendations and parental controls.
Optimizing Communication for Mixture-of-Experts Training with Hybrid Expert Parallel
In LLM training, Expert Parallel (EP) communication for hyperscale mixture-of-experts (MoE) models is challenging. EP communication is essentially all-to-all,... In LLM training, Expert Parallel (EP) communication for hyperscale mixture-of-experts (MoE) models is challenging. EP communication is essentially all-to-all, but due to its dynamics and sparseness (only topk experts per AI token instead of all experts), it’s challenging to implement and optimize. This post details an efficient MoE EP communication solution, Hybrid-EP, and its use in the… Source
Anthropic partners with Allen Institute and Howard Hughes Medical Institute to accelerate scientific discovery
Anthropic partners with Allen Institute and Howard Hughes Medical Institute to accelerate scientific discovery.
Import AI 443: Into the mist: Moltbook, agent ecologies, and the internet in transition
Import AI 443 examines agent ecology systems, Moltbook framework, and adversarial agent corruption risks.
Snowflake and OpenAI partner to bring frontier intelligence to enterprise data
OpenAI and Snowflake announce $200M partnership embedding frontier models and agents directly in Snowflake's data platform.
xAI joins SpaceX
SpaceX acquires xAI, consolidating AI development with rocket/hardware infrastructure.
Introducing the Codex app
OpenAI releases Codex app for macOS enabling parallel multi-agent coding workflows with long-running task support.
Operation “Trolling Stone”: Russia-linked influence activity
OpenAI disrupted Russia-linked influence operation using AI to generate comments on Russian cult leader arrest in Argentina.
Operation “Fish Food”: Russia-origin content farm activity
OpenAI banned Rybar network accounts (likely Russia-origin) using AI for multilingual influence campaigns across websites and social platforms.
Operation “No Bell”: Coordinated criticism of the US and allies
OpenAI banned Operation No Bell, a likely Russia-origin influence campaign using AI to produce anti-US criticism targeting African audiences.
Operation “Date Bait”: AI-enabled scam targeting loveseekers
OpenAI banned Cambodia-origin accounts using AI to conduct romance scams against Indonesian targets via translation and engagement automation.
Operation “False Witness”: Fake recovery service impersonating authorities
OpenAI banned Cambodia-origin accounts using AI to impersonate recovery services and authorities, targeting fraud victims with false recovery schemes.
Romance scams: AI-enabled romance scam workflows
OpenAI banned accounts automating romance scam workflows using AI for outreach, translation, engagement, and investment fraud solicitation.
Silver lining playbook: Likely China-origin activity targeting US persons
OpenAI banned likely China-origin accounts researching US persons and social-engineering tactics to enable targeted influence operations.
“Cyber Special Operations”: China-linked influence planning
OpenAI banned China-linked accounts using AI for influence planning, harassment, and coordinated online operations.
Advancing GPU Programming with the CUDA Tile IR Backend for OpenAI Triton
NVIDIA CUDA Tile is a GPU-based programming model that targets portability for NVIDIA Tensor Cores, unlocking peak GPU performance. One of the great things... NVIDIA CUDA Tile is a GPU-based programming model that targets portability for NVIDIA Tensor Cores, unlocking peak GPU performance. One of the great things about CUDA Tile is that you can build your own DSL on top of it. This post shares the work NVIDIA is doing to integrate CUDA Tile as a backend for OpenAI Triton, an open source Python DSL designed to write DL kernels for GPUs. Source
Establishing a Scalable Sparse Ecosystem with the Universal Sparse Tensor
Sparse tensors are vectors, matrices, and higher-dimensional generalizations with many zeros. They are crucial in various fields such as scientific computing,... Sparse tensors are vectors, matrices, and higher-dimensional generalizations with many zeros. They are crucial in various fields such as scientific computing, signal processing, and deep learning due to their efficiency in storage, computation, and power. Despite their benefits, handling sparse tensors manually or through existing libraries is often cumbersome, error-prone, nonportable… Source
Practical Security Guidance for Sandboxing Agentic Workflows and Managing Execution Risk
AI coding agents enable developers to work faster by streamlining tasks and driving automated, test-driven development. However, they also introduce a... AI coding agents enable developers to work faster by streamlining tasks and driving automated, test-driven development. However, they also introduce a significant, often overlooked, attack surface by running tools from the command line with the same permissions and entitlements as the user, making them computer use agents, with all the risks those entail. The primary threat to these tools is… Source
Project Genie: Experimenting with infinite, interactive worlds
Project Genie lets Google AI Ultra subscribers create and explore infinite interactive worlds via experimental prototype.
Inside OpenAI’s in-house data agent
OpenAI describes internal data agent using GPT-5 and Codex with memory for reasoning over large datasets.
Retiring GPT-4o, GPT-4.1, GPT-4.1 mini, and OpenAI o4-mini in ChatGPT
OpenAI announces retirement of GPT-4o, GPT-4.1, and o4-mini from ChatGPT effective February 13, 2026; API unaffected.
Taisei Corporation shapes the next generation of talent with AI
Taisei Corporation deploys ChatGPT Enterprise for internal HR and talent development workflows.
ServiceNow chooses Claude to power customer apps and increase internal productivity
ServiceNow adopts Claude to power customer-facing apps and boost internal productivity.
Ensuring Balanced GPU Allocation in Kubernetes Clusters with Time-Based Fairshare
NVIDIA Run:ai v2.24 introduces time-based fairshare, a new scheduling mode that brings fair-share scheduling with time awareness for over-quota resources to... NVIDIA Run:ai v2.24 introduces time-based fairshare, a new scheduling mode that brings fair-share scheduling with time awareness for over-quota resources to Kubernetes clusters. This capability, built on the open source KAI Scheduler that powers NVIDIA Run:ai, addresses a long-standing challenge in shared GPU infrastructure. Consider two teams with equal priority sharing a cluster. Source