When Does Model Collapse Occur in Structured Interactive Learning?
Theoretical analysis of model collapse in interactive learning environments where models train on synthetic outputs from other models.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Theoretical analysis of model collapse in interactive learning environments where models train on synthetic outputs from other models.
Comparative study shows structured prompts improve LLM output quality and reduce interaction overhead across ChatGPT, Claude, Grok.
[https://x.com/Google/status/2056789235500466273?s=20](https://x.com/Google/status/2056789235500466273?s=20) Google asked its agents to build a working operating system from scratch using u/Antigravity 2.0 and Gemini 3.5 Flash. Gemini built a real OS out of scratch. It took: ⏱️ 12 hours 🤖 93 parallel sub-agents 🔄 15k+ model requests 🧠 2.6B tokens processed 💸 Less than $1K in API credits To build a functioning OS from scratch.
Quantization benchmark study on Qwen 27B comparing TurboQuant, TCQ, and symmetric q8 via PPL/KLD metrics on RTX 3090 with 64k-128k context.
Goal-oriented calibration method addresses lower-tail miscalibration in Gaussian process Bayesian optimization.
Codegraph tool uses pre-indexed knowledge graphs to reduce Claude API tool calls by 94% and latency by 82% for code analysis tasks.
User requests read-aloud feature parity between Claude iOS app and desktop/web clients.
TrajTok learns transferable trajectory embeddings via adaptive multi-resolution hexagonal spatial tokenization of GPS traces.
FiLark streaming framework enables interactive exploration and real-time processing of distributed acoustic sensing data.
MixRea benchmark tests whether LLMs exhibit inattentional blindness to subtle contextual cues across 2,246 multi-choice questions.
Framework evaluates model-brain alignment by identifying which response dimensions in visual cortex are recovered by neural networks.
Theoretical analysis of computational-statistical tradeoffs for Wasserstein distance estimation from samples.
Lean 4 formalization of IMO 2009 Problem 6 using Aristotle API for AI-assisted theorem proving.
Toto 2.0: open-weights time-series foundation models (4M–2.5B params) achieve SOTA on BOOM, GIFT-Eval, TIME benchmarks.
Andrej Karpathy joins Anthropic as researcher, strengthening AI safety and capability teams.
k-inductive neural barrier certificates for safety verification of partially unknown nonlinear dynamical systems.
Theoretical analysis of representation geometry in JEPAs showing isotropic regularization is suboptimal for structured downstream tasks.
Analytical model of pretraining and linear probing via PCA and linear regression characterizing optimal representation dimensionality.
Hybrid tree construction for speculative decoding reduces memory bandwidth overhead while maintaining verification acceptance rates.
Hugging Face releases Carbon, open DNA foundation models; Carbon-3B matches Evo2-7B SOTA while 275x faster with adapted LLM training for genomic data.
Neurosymbolic framework combining formal argumentation with LLM training for explainable ternary claim verification.
Instance-level shapelets for interpretable time-series classification using learnable temporal discriminative patterns.
ThoughtTrace: 1,058-user dataset pairing multi-turn LLM conversations with self-reported user thoughts and reasoning across 20 models.
Analysis of what evolutionary LLM+search systems actually optimize: algorithmic novelty vs. overfitting to task evaluators in code generation.
BalanceRAG jointly calibrates uncertainty thresholds across LLM-only and RAG branches to reduce unnecessary retrieval while maintaining factuality.
VL-DPO aligns autonomous driving motion forecasting with human preferences using vision-language model guidance and DPO finetuning.
Probability-Conserving Flow Guidance reformulates diffusion guidance through continuity equations to preserve learned manifold geometry under strong conditioning.
CopT reverses chain-of-thought order to draft answers before thinking, reducing token costs when LLMs solve problems without extended reasoning.