Logical Judgments Under Pressure: Diagnosing Syllogistic Stability with Learned Soft Prefixes
Learned soft prefix attacks on syllogistic reasoning expose logical stability limits in Qwen, Gemma models under contextual pressure.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
Learned soft prefix attacks on syllogistic reasoning expose logical stability limits in Qwen, Gemma models under contextual pressure.
Extension of PCMCI+ for causal discovery on irregularly sampled time series (sensor, healthcare, financial data).
One-step and two-step RAG-based policy learning via vector search as nearest-neighbor matching in causal inference framework.
GigaPath-Flash and GigaTIME-Flash: computationally efficient pathology foundation models for whole-slide image and tumor analysis.
SWE-Pruner Pro: prunes long context in coding LLMs by extracting internal relevance signals, improving efficiency over external classifiers.
ATLAS: multi-environment factor model decomposing latent structure into invariant and heterogeneous components for transfer learning.
Context-conditioned safety critic learns adaptive clearance margins for diffusion-based robot navigation in cluttered environments.
PPL-Factory: task-aware perplexity-based data selection for efficient LLM fine-tuning across reasoning and language tasks.
TBSM: one-step generative model using distributional energy and three-body particle interactions instead of adversarial critics.
Certified training approach for convolutional perturbation robustness in vision models with formal safety guarantees.
EVOLVE: autoencoder-based volume compression for scientific simulations with variable-rate encoding.
VEHBench: 763-task diagnostic benchmark for evaluating LLM-assisted vibration energy harvester design across coupled physical constraints.
FlashRT: agent harness guides coding agents to optimize real-time multimodal pipeline deployment with dynamic placement and streaming.
Ben Thompson proposes US law to legalize model distillation and data collection as fair use, addressing licensing hypocrisy and competitiveness vs. Chinese models.
Adaptive Digital Twin framework using robust MPC for continuous surrogate model validation and drift detection in additive manufacturing.
OR Else: smooth one-sided saturation rule replaces PPO clipping for stable LLM post-training with reduced gradient discontinuities.
Mathematical characterization of how temperature scaling distorts soft-label Bayes-error estimators in binary classification tasks.
TRIM reduces verbose AI-generated code by minimizing agent trajectory artifacts through search-process cleanup.
Differentiable Logic Gate Networks enable low-latency EEG classification on edge devices via Boolean circuits instead of floating-point ops.
Neural network feature attribution applied to distinguish totally positive matrices via characteristic polynomial coefficients.
Tutorial on agentic AI architectures for smart grids covering forecasting, optimization, and control with external solvers.
Benchmark evaluating LLMs' ability to reason about 3D spatial constraints in structure-based drug design vs. diffusion models.
O-VAD: training-free agentic framework for industrial video anomaly detection using object-centric tracking and VLM reasoning.
Manifold-Constrained Hyper-Connections enable parameter-efficient finetuning of frozen Transformers via learned residual routing.
ClouDens detects anomalies in high-dimensional cloud system telemetry using context-aware methods for large-scale distributed infrastructure.
Covariance-based surrogate penalty improves subgroup-fair clustering by addressing computational cost of multi-sensitive-attribute fairness.
SGA module detects and fixes geometric errors in LLM-generated pedagogical animation code via symbolic scene graphs.
Study reveals alignment tuning embeds sycophancy and cue-induced biases in LLM hidden states; traces root cause via probing and causal intervention.
Experiential Learning repurposes LLM-as-Judge feedback into coaching signals for policy RL on open-ended tasks, preserving textual feedback bandwidth.
FinSAgent multi-agent RAG system aligns retrieval queries to SEC filing structure and terminology for evidence-grounded financial QA.