AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents
AssayBench: benchmark for LLMs and agents on virtual cell phenotypic screening combining textual inputs with diverse cellular outputs.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
AssayBench: benchmark for LLMs and agents on virtual cell phenotypic screening combining textual inputs with diverse cellular outputs.
Self-Optimizing Language Models (SOL): dynamic per-token compute allocation via lightweight policy network paired with frozen LLM.
CADBench: unified multimodal benchmark for CAD program generation with 18k samples across six modalities and design datasets.
Attractor-Vascular Coupling Theory: mathematical framework for cuffless blood pressure estimation from smartphone photoplethysmography.
Decision-centric rate-distortion framework for agent memory compression prioritizing decision quality over descriptive faithfulness.
Qwen3.5 0.8B sees 2.88M monthly downloads; user reports semantic understanding, JSON parsing, and latency challenges in production workflows.
BEACON: 430GB multimodal dataset of Valorant gameplay for behavioral authentication and continuous monitoring across skill tiers.
BenchCAD: comprehensive benchmark for programmatic CAD generation from visual/textual inputs in realistic industrial settings.
Directional Groupwise Preference Optimization (DGPO): group-level margin-based framework for LLM alignment with directional consistency.
RUBEN uses rule extraction and pruning to explain RAG-LLM outputs and test safety robustness against adversarial prompts.
MiniCPM 4.6 released on Hugging Face; open-weights efficient model variant with updated capabilities.
EditMGT applies Masked Generative Transformers to localized image editing, outperforming diffusion-based approaches.
Digg returns (again) as another place to read AI news.
Counterfactual data augmentation improves Vision-Language Models' chart understanding efficiency without scaling synthetic datasets.
RAG-based satirical definition generator for Finnish news context with human-annotated evaluation framework.
Generalized Turing Test formalizes agent intelligence comparison via indistinguishability, independent of tasks or datasets.
Pi-Serini evaluates BM25 lexical retrieval sufficiency in agentic research loops paired with frontier LLMs.
Distance-metric-based instance methods detect conditional anomalies in patient management alerts.
Reddit user compares leaked Gemini Omni video model against Sora 2, which OpenAI is reportedly discontinuing.
BabelDOC preserves PDF layout during cross-lingual translation via intermediate representation decoupling structure from text.
DISCA steers LLM cultural preferences via sociodemographic disagreement signals without fine-tuning or white-box access.
Clin-JEPA extends joint-embedding predictive pretraining to EHR trajectories for multi-task patient risk prediction.
Transcoda applies synthetic data and Humdrum kern encoding to optical music recognition without large labeled datasets.
Pentesting agents evaluated on real-world targets show current benchmarks miss complexity and strategic decision-making required in practice.
MMVIAD introduces first multi-view video dataset for industrial anomaly detection with continuous 2-second inspection clips.
Framework enables visual-native multimodal search agents with on-policy data evolution and persistent visual evidence reuse.
SLIM uses sparse autoencoders to steer LLM hidden states for interpretable and controllable molecular property editing.
Combines NeRF and diffusion models for probabilistic 3D scene reconstruction via latent posterior sampling.
Long-context LLM performance degrades nonlinearly with misleading information proportion, critical for RAG and agentic systems.