Entity-Faithful Repair of Synthetic Supervision for Zero-Shot Image Captioning
Method to repair synthetic training data for zero-shot image captioning by fixing entity misalignment at fine-grained level.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Method to repair synthetic training data for zero-shot image captioning by fixing entity misalignment at fine-grained level.
SCHEDBench: benchmark with 1,132 scheduling instances to evaluate LLM constraint faithfulness under natural-language variation.
FDR-controlled feature selection framework for grouped features in sequential and neural models using block-level mirror statistics.
Contrastive pretraining framework for single-cell transcriptomic foundation models to learn cell representations beyond gene reconstruction.
Transfer learning approaches for Named Entity Recognition on small, unlabelled datasets across multiple domains.
Microsoft-led open letter signed by 235 AI companies including NVIDIA and OpenAI argues against US government restrictions on open-weight models on safety grounds.
Two-sided audit framework for self-improving AI-for-science systems to distinguish real gains from search artifacts and oracle drift.
Simon Willison's June 2026 newsletter roundup covering model releases (GPT-5.6, Claude Opus 5, DeepSeek-V4), open letters, and accidental cyberattacks by OpenAI and Anthropic test models.
Study of LLM persona panels (GPT-4.1) as synthetic research participants, showing marginal-check validity depends on prompt and uncertainty assumptions.
Quantile Coupling Flow Matching: lightweight coupling method for flow-matching generative models with subquadratic cost in batch size.
Search-GRT: RL method to train LLM search agents for multi-hop QA with guided retrieval to reduce sparse-reward training issues.
Zero-query jailbreaks for text-to-image systems via filter-generator discrepancy; transfer-based attack requiring no target queries.
PROGRESS trains search-augmented LLM agents using coverage-guided RL rewards to improve query decomposition over outcome-only supervision.
TrajWiki proposes trajectory-based external memory for long-horizon dialogue agents with traceable, updatable, diagnostically transparent storage.
CryptoProver synthesizes proofs for cryptographic libraries by inferring internal specifications from API contracts, verified on curve25519-dalek and chacha20.
PMMC compiles multimodal memory at consolidation time for LVLM agents to preserve image-text binding and temporal updates without query-time overhead.
Essay critiquing institutional dismissal of AI anthropomorphism as user error, arguing for epistemic pluralism in human-AI interaction interpretation.
SHAP-based interpretable ML framework for predicting asphalt concrete splitting strength using TabPFN, XGBoost, and classical ML baselines.
Data-driven elastic-net SVM with learnable simplex-constrained weights over candidate pinball losses, with empirical oracle inequality bounds.
GraRe re-ranks 6-DoF grasp candidates using geometry and object context without modifying frozen detectors, improving alignment with actual grasp quality.
Benchmarking PPG-based sleep staging datasets and metrics, showing substantial performance gap vs. EEG and advocating finer temporal resolution.
Analysis of entropy inversion in LLM temperature scaling: autoregressive feedback drives 11 models through entropy maximum into population inversion and nonlinear dynamics.
Neuro-symbolic governance framework for verifiable AI agents in decentralized digital twin ecosystems with semantic profile layers.
xMICD balances interpretability and performance in ICD code representation for clinical risk prediction via embedding-based explainability.
Gaokerena: compact Persian-language medical LLM family trained on 90M-token corpus for low-resource healthcare deployment.
LAND model simulates 314K heterogeneous agents and human actors over 30 days to study emergent social dynamics via LLM-enabled ABM.
Calibration breaks under unseen subtype shift: models maintain accuracy but become systematically overconfident on novel fine-grained categories.
RefactorAssist agentic system improves LLM code refactoring reliability by detecting and correcting functional behavior changes pre-deployment.
Tevatron 3.0 integrates Megatron-Core for efficient MoE reranker training, enabling billion-scale cross-encoder + distillation workflows on academic budgets.
UpliftBench reveals metric disagreement, not model disagreement, drives uplift estimator ranking variance across 7 dataset families and 12 methods.