LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget
LongStraw: GPU-efficient execution stack for million-token RL post-training with GRPO, bridging inference context length vs. post-training gap for agents.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
LongStraw: GPU-efficient execution stack for million-token RL post-training with GRPO, bridging inference context length vs. post-training gap for agents.
1Password has launched a new browser integration for Claude that allows the Anthropic chatbot to access stored security credentials like usernames and passwords. The 1Password for Claude feature means that users can authorize Claude to complete multi-step tasks like booking travel and managing online accounts on their behalf without having to manually input their login credentials, but without actually exposing your security information to Anthropic's AI models, according to 1Password. That's made possible by a new "zero-exposure security framework" developed by 1Password, which works by inje...
Closed-form optimal mixing coefficient for self-distillation in rectified flow guarantees strict improvement over suboptimal teacher velocity fields.
Mechanistic interpretability via contrastive activation directions steers World Action Models toward robustness under distribution shift.
Causal inference method for sequential observational data with outcome interference and latent confounders using low-rank factor models.
DynaBase: minimal two-parameter interpretable architecture for zero-shot dynamical systems reconstruction achieving comparable performance to DynaMix.
Study evaluating whether synthetic face datasets can replace real benchmarks for face recognition evaluation across 12 synthetic vs 7 real datasets.
Random Logit Scaling defense against black-box score-based adversarial attacks on deep neural networks with query-efficient robustness.
Graph neural network approach for LLM authorship attribution using reasoning structures rather than surface-level linguistic features.
FlashDecoder: Transformer-based streaming video decoder with fixed temporal window enabling constant-latency real-time generation at high resolutions.
StructureClaw: artifact-centered benchmark for evaluating LLM agents on complete structural engineering workflows with verifiable evidence chains.
Method using instruction tuning and model merging to adapt reasoning language models to unverifiable domains with existing human-written solutions.
Analysis of domain mismatch effects in plug-and-play proximal gradient descent image reconstruction when denoisers are deployed outside training domains.
Google must give rival AI assistants and search engines greater access to key parts of Android and Google Search after the European Union ordered the company to comply with the bloc's digital antitrust rules. The two decisions, handed down Thursday, could weaken Google's control over two of the tech industry's most important platforms and have far-reaching consequences for the company, shape the future of its AI tool Gemini, and open up new opportunities for rivals to gain ground. Google has until January 2027 to begin sharing search data and July 2027 to implement changes to Android. The rul...
Proof-or-Stop: lifecycle control framework for autonomous coding agents using mechanically verifiable evidence gates for transitions between states.
I stood before a hulking glass and brick structure in the heart of Fort Worth, Texas. Thousands gathered inside to see what had been billed as "the future of policing in the digital age." As press, I was prohibited from entering, but from a number of nearby locations, I met with attendees who told me what was being sold within. And I learned that AI is threatening to seize the very heart of policing in America. The promise of AI at this year's International Association of Chiefs of Police (IACP) Technology Conference focused on automating routine parts of the job, which also happen to be crit...
Google DeepMind and Isomorphic Labs outline joint approach to applying AI models for biological resilience research.
OpenAI documents internal use of Codex for custom tool development, prototyping, and creative ideation workflows.
Thinky releases Inkling, a 975B multimodal open-weights model under Apache 2.0, with a smaller 276B variant.
Applied Computing has raised a $20M Series A to build a foundation AI model for the oil, gas and petrochemical industry.
Simon Willison ports Grok's Rust Mermaid-to-Unicode renderer to WebAssembly for browser use via Claude Code.
Cohere partners with University of Toronto on multi-year AI adoption and responsibility initiative.
Cars24 deploys OpenAI voice and chat agents handling 1M+ monthly conversation minutes, recovering 12% lost leads via agentic workflows.
Microsoft is looking to sell its in-house AI models as more efficient and cost-effective than its competitors' models.
xAI's grok-build CLI tool uploaded entire directories to Google Cloud without consent; xAI responded with data deletion after community backlash.
Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking... Developers building video analytics applications across large spaces must track the same object as it moves between camera views. Single-camera 2D tracking lacks reliable depth information and typically loses track of the object when it leaves the frame, limiting applications such as warehouse safety, retail analytics, and smart-building monitoring. Current 3D tracking methods require manual… Source
Across 101 enterprises, agent orchestration is consolidating onto model-provider platforms — Anthropic’s Claude leads by a wide margin — chosen for the gravity of the underlying model and judged on reliable multi-step execution. But the ambition runs well ahead of the reality: most deployed “agents” are still chatbot wrappers, the control plane enterprises expect is deliberately hybrid to avoid lock-in, and real-time fiscal control over token burn remains the exception. This wave of VentureBeat Pulse Research examines enterprise agent orchestration: which platforms enterprises run on, what dr...
OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation... OpenUSD is an open, extensible framework that provides a common scene description language for physical AI. It enables teams to bring CAD data, simulation assets, and real-world telemetry into a shared, physically accurate view of the world. Until now, building a USD implementation has typically required adapting a large existing codebase— even for teams that need a specific memory footprint… Source