Visual Credit Audit for Multimodal Spatial Reasoning
Visual Credit Audit (VCA) measures whether multimodal models genuinely use image information vs. text-only reasoning in spatial benchmarks.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Visual Credit Audit (VCA) measures whether multimodal models genuinely use image information vs. text-only reasoning in spatial benchmarks.
SciFigAlign evaluates scientific figures by alignment of visuals with manuscript claims, addressing peer-review assessment beyond generic image quality.
ScratchSim generates synthetic industrial defect data via procedural rendering (BlenderProc) with domain randomization for surface scratch detection.
PIKS develops kernel methods for physics-informed learning with closed-form solutions and analytical theory, contrasting with PINN complexity.
Microsoft is on a mad dash behind the scenes to patch exploits before hackers find them.
Setoka benchmark evaluates hierarchical user understanding in memory-augmented personalized agents beyond explicit fact retrieval.
CoCaRS uses correlation calibration to suppress redundancy in heterogeneous knowledge distillation across diverse model architectures.
AI home management startup Hint, co-founded by Martha Stewart, wants to become an “AI for your home,” combining property records, maintenance schedules, home documents, and an AI assistant into a single app.
GPTQ-2D extends adaptive matrix rounding (GPTQ) to two-sided case with nonsingular basis matrices for quantization optimization.
Video world models suffer dimensional collapse during long autoregressive generation; representation regularization stabilizes frame quality over 100+ steps.
Sparse lottery tickets matched to dense model accuracy fail to maintain behavioral equivalence in production deployment scenarios.
HoF-Bench evaluates 95 real CVEs discovered by LLM analyzers (OpenSSL, curl, GnuTLS) without frontier models, establishing reproducibility bar for AI security scanning.
TreeCCA applies gradient-boosted trees to canonical correlation analysis via Eckart-Young loss for tabular feature extraction.
BayesAME automatically determines coreset size for efficient LLM benchmark evaluation using sequential Bayesian inference.
Stereotypes-to-Decisions framework measures regional bias in six LLMs across China's 34 provinces, linking abstract stereotypes to concrete allocation decisions.
Certificate-gated protocol determines which physical parameters (mass, drag, stiffness) latent world models actually internalize from raw visual observations.
OpenAI reports 3x improvement on ARC-AGI-3 benchmark via two API settings enabling reasoning retention and compaction in GPT-5.6.
Overlap gap property extended to finite-temperature neural network optimization, characterizing algorithmic accessibility of noisy solutions.
AgentSnare uses adaptive deceptive observations to mislead LLM-based penetration testing agents, defeating static artifact recognition.
24 frozen vision foundation models (ViT, CLIP, etc.) evaluated for face presentation attack detection via linear probing under domain shift.
Algebraic framework deriving transformer expressivity from finite-precision attention dynamics and memory constraints.
The startup analyzes calls, messages and CRM data to identify effective sales techniques and turn them into playbooks for AI agents.
SymmGrid framework accelerates on-robot reinforcement learning via parallelized symmetries and dual visual perception.
LLM-based validation of atom-centered structural descriptors reveals descriptor degeneracies in materials science.
OptimismBench detects directional bias in LLM probability judgments via inverted-pair evaluation method.
TREK benchmark evaluates LLM agents on complex, executable travel itinerary planning with verifiable constraints.
Pair-level judgment outperforms dialogue-level generation for emotion-cause extraction in conversation datasets.
Feature stability analysis framework shows feature bagging ensemble reduces generalization error.
BAND sparse Bayesian network achieves polynomial convergence rates for high-dimensional distribution estimation.
CreditCardQA benchmark evaluates LLM numerical reasoning on real financial documents with CoT/PoT comparison.