AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents
AgenticSTS introduces bounded-memory architecture for LLM agents with typed retrieval instead of transcript accumulation for long-horizon tasks.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
AgenticSTS introduces bounded-memory architecture for LLM agents with typed retrieval instead of transcript accumulation for long-horizon tasks.
Theoretical proof that aggregation with exponential weights achieves minimax-optimal excess risk in expectation for model selection.
Copewell deploys multi-agent swarm for mental wellness support in low-resource settings with dynamic emotional state calibration.
FTC urged to reject Elon Musk’s bid to end X monitoring amid AI concerns.
Paul Bakaus on Impeccable discusses skill engineering, human-in-the-loop agent design, and limitations of one-shot prompting approaches.
Survey identifies limitations of LLM-as-a-Judge evaluation in multilingual and low-resource language settings with inadequate validation.
Purified OPSD fixes on-policy self-distillation failures on long chain-of-thought by decomposing teacher supervision signals.
Confidence-guided classification model for AI-assisted household waste sorting across German municipalities using human-in-the-loop.
CoFL-S predicts language-conditioned flow fields for low-level robot navigation actions in vision-language environments.
SpeechCombine enables speech language models to follow instructions without explicit instruction tuning via text-speech knowledge transfer.
MLP-GNN framework isolates chemical and structural contributions to aqueous solubility predictions for drug discovery.
Guard Rail Validation framework standardizes runtime interception and validation of AI agent decisions in autonomous telecom networks before execution.
Conformal prediction framework for building coverage-guaranteed prediction sets in counterfactual decision-making with uncertainty quantification.
Self-explainable operator learning framework reformulates neural architectures as interpretable linear combinations of functional models for physics systems.
Eticas AI Risk Taxonomy operationalizes AI audits by bridging 74+ competing risk taxonomies with executable tests, measurements, and defensible severity grades.
H-Score objective enables mutual information-inspired feature learning with proof of invariance to invertible transformations under constrained approximation.
Taxonomy of 53 human-AI team studies reveals five clusters (AI Assistant, Ad-hoc Dependency, Paired Equanimity, Group Equanimity) with distinct characteristics.
Overview of AI risk assessment and management methodologies across regulatory frameworks including EU AI Act with techniques for technical and ethical risks.
Online resource allocation algorithm bounds regret for sequential requests with continuous reward/consumption distributions and degenerate fluid relaxation.
DSGNAR optimization framework solves ill-conditioning in physics-informed neural networks via doubly-sketched Gauss-Newton with adaptive ratio.
Model-agnostic framework for privacy-preserving distributed computing unifies federated and decentralized learning defenses against adversarial manipulation.
UA-ChatDev mitigates hallucination propagation in LLM-based multi-agent software development by tracking uncertainty across collaboration stages.
RadiomicNet combines handcrafted radiomics features with deep learning for interpretable medical image segmentation.
Microsoft follows Amazon, OpenAI and Anthropic with its new AI deployment group.
DALorRA applies Bayesian sparse low-rank adaptation to quantify LLM uncertainty and reduce overconfidence in fine-tuned models.
Rubric-based clinical reasoning benchmark shows frontier LLMs achieve ~32% on hard expert-authored medical cases.
Dynamic Neural Graph Encoding uses temporal graphs to analyze neural network weight spaces and inference sequences.
Proves multi-secretary problem requires Ω((log T)²) regret for bounded-density distributions with support gaps.
Ensemble machine learning methods detect early-stage Alzheimer's disease biomarkers from neural network analysis.
A²utoLPBench auto-generates infinite linear programming word problems via inverse-KKT for LLM-agent evaluation.