TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
TACT: taxonomy-aligned post-training framework for pedagogically adaptive LLM-based ESL tutoring with human-grounded evaluation.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
TACT: taxonomy-aligned post-training framework for pedagogically adaptive LLM-based ESL tutoring with human-grounded evaluation.
llm 0.32 release includes updates enabling new Claude model support and streaming typed events for reasoning and tool results.
Muon optimizer shows token-efficiency gains on Mamba-2 when applied selectively to output projections, with benefits localized to specific weight groups.
Logic pre-pretraining on formal derivations accelerates LM skill acquisition and improves compressibility compared to narrow symbolic tasks.
Mariano-Florentino Cuéllar joins Anthropic as first Chief Global Affairs Officer, signaling focus on policy and governance.
Latent Reward Registers enable preference alignment in Diffusion Transformers via learnable tokens that extract reward signals from intermediate noisy latents.
R-ItCUR algorithm recovers low-tubal-rank tensors from partial cross-concentrated sampling observations with sparse outlier corruption.
Physics-Flavored Neural Network automates kinetic phenotyping of engineered skeletal muscle tissues, replacing simplistic peak-force metrics.
PRISM meta-workflow converts multivariate time series to multi-channel images for anomaly detection, evaluating vision backbone alternatives.
Transformers execute Sequence-level Interactive Dynamic Parallel Processing (SIDPP), constructing prompt-dependent parameter transformations at inference rather than reproducing training statistics.
Standard music transformers lose equivariance to pitch transposition and time shifts as scale increases, allocating capacity to memorizing absolute patterns.
EcoFrame enables training-free adaptive frame scheduling for long video understanding in VLMs, using inference feedback to adjust sampling budgets.
First implementation of causal perception framework for competing Structural Causal Models, operationalizing fairness in agent reasoning systems.
Acceleration Matching algorithm for smooth trajectory inference from discrete snapshots.
Sparse Weight Decomposition enables efficient mechanistic interpretability of pretrained transformers without sparse retraining.
Social theory framework for pluralistic AI alignment recognizing multiple legitimate perspectives in diverse deployment contexts.
Lipschitz singularities in dynamic routing transformers create adversarial vulnerabilities in efficient UAV tracking systems.
ANNOTARES dataset for extracting logical structures (conditions and consequences) from German statutory texts.
KV cache transfer enables prefill reuse across different-sized LLM family models via closed-form linear mapping.
Contrastive activation addition steers temporal horizon preferences in Qwen3-32B via learned linear representations.
CARE-X unifies chest X-ray report generation with classification, localization, and measurement for clinical VLM utility.
Omega-S penalty retains fine-tuned model knowledge without prior-task data, Fisher matrix, or weight copies.
BanglaWild benchmark evaluates 15 VLMs and OCR systems on 2,535 in-the-wild Bengali scene text images.
Zero-shot LLM reranking pipeline for ADHD symptom sentence ranking from Reddit text using BM25, embeddings, and semantic rescoring.
MultiGlobeQA: 46K-example multilingual benchmark for geospatial reasoning exposing LLM failures in spatial relations across 14 function families.
Feasibility-aware generative model for AC-operable synthetic power-grid scenarios for resilience and contingency analysis.
Structure-Aware Fine-Tuning (SAFT) refines VLM reward models via self-supervised learning without ground-truth labels for RL.
ContinualSkillBench evaluates whether LLM agents can continually acquire and reuse skills across 500 interconnected task domains.
GENESIS: Explainable causal discovery from observational data combining LLM reasoning with statistical methods to justify edge decisions in DAGs.
ADMITBench: Safety-governed evaluation framework for industrial LLM recommendations assessing evidence support, authority, and plant-specific consequences.