OpenForgeRL: Train Harness-native Agents in Any Environment
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.
VCSD proposes visual contrastive self-distillation removing need for privileged information in on-policy distillation via pure input conditioning.
MIRROR framework exploits complementary reasoning paths across text, diagram, and combined modalities to improve vision-language model reasoning on geometry problems.
X³-OPD cross-modal distillation framework transfers reasoning from text LLM teacher to audio-language student via on-policy alignment and acoustic perception.
Neural networks solve coupled Dyson-Schwinger equations for Yang-Mills gauge theory with percent-level agreement to fixed-point solutions.
Theory paper argues human participation persists in automated systems for technical, complementarity, and normative reasons beyond current AI capability limits.
Zero-Flow Two-Sample Test uses learned directional misalignment patterns for distribution testing, separating witness learning from hypothesis evaluation.
DONDO releases 26 open w2v-BERT speech recognition models for African languages spanning six countries, trained on religious text corpora.
Windowed-MTP optimizes speculative decoding at million-token context by eliminating full-KV attention overhead in multi-token prediction draft heads.
Petri-net-guided LLM test generation for concurrent Rust APIs addresses shallow test synthesis by integrating formal models with executable test concretization.
ElasticTTT framework prevents prior collapse in test-time tuning of diffusion models for video editing by preserving distribution-mapping during optimization.
Runway no longer wants to be just another AI model company. It wants to become the infrastructure layer for generative media. On Thursday, the startup launched Runway Media Router through Runway Dev, its developer platform, released earlier this month, that provides API access to a growing roster of third-party image, video and audio models alongside […]
GS-Agent generates physically plausible 4D worlds from natural language by combining foundation models with agentic simulation and physics constraints.
Study using gpt-5.6-sol shows LLMs produce safer advice when dangerous objectives are mediated through agent transformation versus direct exposure.
Improved lower bounds for Shannon capacity of odd cycles via independent set construction in graph powers—pure graph theory unrelated to AI.
Users can also integrate their personal data from services like Apple Health, Function, and MyFitnessPal.
OpenAI is rolling out ChatGPT Health to everyone in the US on Thursday, allowing more people to connect their medical records and health-tracking information to the chatbot. During a briefing, Ashley Alexander, OpenAI's vice president of health product, says the company's models "are now capable of reasoning at levels that are better than clinician level." When asked for more information about how the performance of OpenAI's models stacks up against human clinicians, OpenAI health lead Karan Singhal says that he would "temper" the claim that they are reasoning at better levels, but that "ther...
Agentic context management frames token cost and memory bloat as lifecycle and architecture problems, not storage-retrieval, for production agent reliability.
LLMs systematically overuse epanorthosis (classical self-correction rhetoric) due to promotional training distributions and RLHF preference for emphatic phrasing.
Speech-based multimodal LLMs detect cognitive impairment across diverse speakers and devices by leveraging linguistic and acoustic biomarkers with improved generalization.
No-code agent platforms create reliability gaps—silent degradation from changing models, tools, permissions, and dependencies—requiring continuous assurance frameworks.
Analysis of code model representations shows Qwen2.5-Coder and DeepSeek-Coder align on grammatical concepts across Python/Rust, with task-driven specialization.
David Bowie's song "Five Years," which Meta used in a supposedly inspiring advertisement, is about humans learning that they have five years left to live before the apocalypse.
MAPS: hierarchical MARL system using centralized proto-plan embeddings for decentralized AV coordination at unsignalized intersections.
Open-source evaluation framework for open-weight LLM agents on longitudinal data tasks, addressing privacy constraints in research deployments.
Label complexity bounds for auditing high-recall candidate generation pipelines with finite-sample validity guarantees.
Randomized KV-cache eviction with error certification via Hájek correction, proving deterministic eviction hides information loss.
Thinkink: 2D spatial interface integrating handwritten/sketch prompts with LLM responses via semantic tree interpretation.
NVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways... NVIDIA OptiX ray tracing engine is an application framework for achieving optimal ray tracing performance on the GPU. Applications using OptiX can fail in ways that are difficult to diagnose: an invalid API argument, a black frame, or a GPU-side bug buried under thousands of concurrent threads. Debugging facilities in the NVIDIA OptiX Toolkit (OTK) can help. OTK is a GitHub repository… Source
AREX: recursively self-improving research agent exploiting discovery-verification asymmetry to refine multi-constraint answers.