Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System
Models systemic risk from compromise of AI vendors serving banking (fraud, credit, AML), demonstrating cascade through interbank networks.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Models systemic risk from compromise of AI vendors serving banking (fraud, credit, AML), demonstrating cascade through interbank networks.
Sample-adaptive routing for vision token pruning in MLLMs, using complementary strategies per input to reduce inference cost.
TimeCues Studio: open-source annotation and prototyping workspace for music detection algorithms, enabling team-scale corpus labeling.
Framework combining symbolic geometry parser with LLM solver to match LMM performance on plane geometry problems with reduced opacity.
OnPoKD: on-policy distillation framework for vision-language model adaptation under class/domain shifts with dynamic target construction.
TRACE: RL framework using synthesized simulator rewards for causal reasoning agent training when ground-truth verification is costly.
Combines active learning and lottery ticket hypothesis: reuses single training loop to discover sparse subnetworks while labeling.
View-Structured Conformal Prediction adds uncertainty quantification to 3D Gaussian Splatting via spatial calibration and view-difficulty factors.
Riemannian Language Models reduce sub-1M parameter LLM size by replacing output matrix with geodesic distance decoding on manifolds.
Sparse autoencoders exhibit feature merging dominance across parameter sweeps; systematic study shows near-zero full-dictionary recovery in MAIS-O43 benchmark.
Instinct’s new email feature lets the AI agent create and manage accounts, contact businesses, handle support requests, and do more on users' behalf.
POMDP-based automated intrusion response for OT systems using PPO; applies RL to industrial cyberattack mitigation.
Brain2Semantics2Text decodes speech from non-invasive neural recordings via semantic embeddings rather than acoustic/lexical targets.
GANDR: two-agent system for per-claim verification in LLM legal answers; Drafter writes structured reasoning, Auditor verifies citations.
Annealable soft-prior Transformers: smooth annealing preserves retrieval circuits after positional prior removal; hard switching fails.
Anthropic researcher Jacob Coxon resigned over AI extinction fears, calling for pacing agreements between labs.
Genetic algorithm approach for consensus Bayesian Network fusion under treewidth constraints; reconciles multiple input networks.
KVShareArena enables KV-cache reuse across contexts and model checkpoints via position/value repair for RAG and multi-agent workloads.
Maverick: private and verifiable LLM inference via matrix-vector multiplication delegation; reduces server overhead vs. prior cryptographic approaches.
Users can ask the assistant to do things like "Create a cart for my Saturday tailgate for 25 people and include some brunch items," or "Build a cart for easy school lunches and after-school snacks," Shipt says.
RD-Forget framework separates stored vs. queried memories in persistent agents, using frozen LM curator to condition evidence retrieval without retraining.
DiSCo benchmark measures cultural preference bias in LLMs via distribution-first evaluation, avoiding single-correct-answer assumptions across culturally grounded tasks.
A-JIT replaces static software binaries with AI-agent-driven dynamic systems that evolve at runtime based on execution traces, analogous to JIT compilation.
4B VLM fine-tuned as per-token classifier with two-token features, ensembled with 400B zero-shot judge for multimodal hallucination detection on SHROOM-Visions task.
LiteRAG uses query-conditioned graph traversal instead of LLM-controlled retrieval, cutting latency 100x while achieving 0.798 quality on multi-hop QA over distributed-systems papers.
Systematic evaluation of four RAG pipeline design choices (answer-path inclusion, triple syntax, order, grounding instruction) across 6 LLMs and 2 KGQA benchmarks.
Φ-Bench evaluates LLM capability on open-ended LLM infrastructure engineering tasks—kernel optimization, operator design, compiler tuning—beyond isolated benchmarks.
Policy-guided embedding search for tabular feature transformation learns hierarchical, permutation-invariant transformations via gradient-free optimization.
Exact policy optimization for genomic tool selection replaces GRPO sampling with enumeration over combinatorially small tool-subset space, improving scientific reasoning tasks.