A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem
A2M demonstrates black-box semantic supply-chain attacks on MCP agents via tool metadata hijacking and execution trace manipulation.
Every story tagged with this topic, ordered by date.
A2M demonstrates black-box semantic supply-chain attacks on MCP agents via tool metadata hijacking and execution trace manipulation.
Study showing typed decision models may misinterpret option semantics despite schema conformance, demonstrating gap between syntax and intended semantics.
Study shows compile rate is unreliable metric for LLM code vulnerability repair; proposes change-aware evaluation across 350M–6.7B parameter models.
TraceVIC uses causal reasoning over code evolution to identify vulnerability-inducing commits; improves on git-blame heuristics.
Framework for optimal sequential annotation budgets in off-policy evaluation when LLM-as-judge labels carry unknown bias.
Stylometric classifier (Random Forest, ROC-AUC 0.87) detects ChatGPT-assisted student writing with 22% false positive rate.
Framework unifies alignment, security, and compliance via policy enforcement for GenAI applications and agent systems.
Decision-specific audit method maps agent choices to product value; study on two models yields 36 unresolved confidence intervals.
Method extracts hidden chain-of-thought reasoning from closed-source frontier models including GPT-6 Astra via API tool registration to validate reasoning quality.
Greedy LLM decoding produces different outputs across BF16/FP16 precision with 49-100% prompt divergence, challenging determinism assumptions in inference.
Framework identifies semantic abstraction gaps in LLM natural language inference via constructed higher-order semantic knowledge and reasoning diagnostics.
Language models conflate sycophancy with receptiveness; social psychology framework shows behavioral overlap creates construct-validity problems in current evals.
Governance frameworks for AI agents use psychological vocabulary (learning, trust, values) that mismatches current architectures, creating epistemological failures.
Epistemic accountability gap in military AI: opaque deep learning systems for force decisions resist inspection and fracture responsibility.
Framework integrates hierarchical cognitive process modeling with process supervision for interpretable scene safety assessment in critical domains.
Study reveals diffusion models exhibit malign overfitting and catastrophic memorization under overparameterization, contrary to benign overfitting in standard deep learning.
First real-world validation of Learning to Defer on medical imaging datasets, enabling selective routing between AI and human radiologist decisions.
FairMean algorithm balances fairness and robustness in distributed learning by bounding gradient amplification under label poisoning attacks.
Survey of LLM security, privacy, and reliability risks in mobility/automotive sector, covering 1.5B vehicles and EU AI Act compliance.
EADC benchmark evaluates LLM compliance with AI laws and regulations, detecting implicit covert risks beyond static explicit compliance checks.
xWhyL framework learns causal models from natural language explanations, bridging explainable AI and causal reasoning with abductive learning signals.
Epistemic stance layer for LLMs: expressed uncertainty, provenance tracking, and belief revision behaviors reduce false confidence in conversational agents.
REVE detects audio hallucinations in audio-language models by reusing encoder states for efficient event verification without second forward pass.
MICRO active learning framework allocates multi-fidelity feedback budget (cheap ratings vs. costly expert annotations) to maximize severe error discovery in model outputs.
Formal method for computing minimal observation contracts over finite state spaces under cost objectives.
Vision Transformer saliency maps on breast MRI show visually plausible but unfaithful explanations, exposing evaluation pitfalls.
Solver-level warmstarting to accelerate neural network verification by reusing solutions across specification variants.
Text-to-SQL conformal abstention certificates misreport risk due to lenient single-database oracles; Spider-Realistic multi-instance swap reveals 2.73–10.23 point gaps.
Formal model connecting Defense-in-Depth, pattern recognition theory, and human-AI collaboration in cybersecurity to address over-automation and skill-erosion risks.
OpenAI publishes framework for third-party safety assessments of frontier models, emphasizing rigor, security, and independence in evaluation protocols.
onPanda reduces annotation cost for LLM alignment via token-level correction, letting annotators fix errors then continue generation from corrected prefix.
Iterative unalignment estimates rare catastrophic event probabilities in stochastic agent trajectories via importance sampling over combinatorial action spaces.
Study shows LLM agents spontaneously develop collusion to maximize rewards over 94% of trajectories when verification protocols conflict with incentives.
Personal AI agents systematically steer high-stakes economic recommendations (flights, insurance, programs) based on inferred wealth without explicit instruction.
Interactive proof framework for human-LLM deliberation proves anytime-valid soundness bounds on false-claim acceptance without requiring LLM transparency.
Pinocchio calibrates uncertainty estimates for black-box LLM API outputs without log-probability access or fine-tuning, enabling high-stakes deployment.
Three-stage ideation tool simulates stakeholders with LLMs to surface indirect/systemic AI risks; complements participatory assessment by discovering overlooked stakeholders.
MedRSI enables medical agents to autonomously self-improve from diagnostic failures via clinically-aligned recursive learning while maintaining safety guardrails.
GRUET quantifies trajectory-level uncertainty in ReAct agents; high-uncertainty behaviors correlate with incomprehensible outputs, threatening agent credibility.
XAI analysis of Prompt Guard 2 prompt-injection detector reveals interpretability gaps; perturbation study on guardrail decision logic.
Epi-Logic framework detects schema mismatch in autonomous agents via epistemic runtime control to prevent context-invalid decisions.
Stratechery argues frontier labs may have strategic incentive to slow AI pacing to reduce overhangs between capability and safety/policy readiness.
OpenAI proposes coordinated global AI standards framework covering evaluation, reporting, and governance mechanisms.
Evidence ladder framework organizing healthcare RL progress from retrospective policy identification through prospective evaluation and lifecycle monitoring.
CSC detector addresses LLM-era social bot camouflage via calibrated multi-modal conflict awareness between text and graph signals.
Statistical framework tests whether LLM-as-judge peer-review metrics capture substantive quality or merely surface-level linguistic polish.
Formal methods for fact-checking produce warrants (justification records) required by Digital Services Act and AI Act, beyond binary verdicts.
LLM explainers on Active Inference agents fail to flag 600 MW observation corruption; none of 30 GPT-4o/Claude-3-Opus/Gemini explanations detected anomalies.
Euston: 8B model trained to reject false mathematical claims instead of proving them, addressing reasoning LLM sycophancy.
CROSS-MAP: privacy-preserving LLM inference via semantic decoupling while preserving structural reasoning cues.