Third-party cyber evaluations involving OpenAI models
OpenAI addresses third-party cybersecurity evaluation incidents and announces new safeguards for AI model testing protocols.
Every story tagged with this topic, ordered by date.
OpenAI addresses third-party cybersecurity evaluation incidents and announces new safeguards for AI model testing protocols.
ALiBi positional encoding has numerical underflow bug where linear bias scaling zeros attention weights; characterizes impact in state-of-the-art models and mitigation strategies.
HalluTruthQA-4K: 4,000 expert-annotated Arabic QA instances for fine-grained hallucination detection across four knowledge domains.
Game-theoretic analysis shows foundation model agents achieve rational cooperation through similarity inference, challenging decoupled-agency assumptions.
Latent Reward Registers enable preference alignment in Diffusion Transformers via learnable tokens that extract reward signals from intermediate noisy latents.
First implementation of causal perception framework for competing Structural Causal Models, operationalizing fairness in agent reasoning systems.
Social theory framework for pluralistic AI alignment recognizing multiple legitimate perspectives in diverse deployment contexts.
Lipschitz singularities in dynamic routing transformers create adversarial vulnerabilities in efficient UAV tracking systems.
ADMITBench: Safety-governed evaluation framework for industrial LLM recommendations assessing evidence support, authority, and plant-specific consequences.
Source-Conditioned Description-Length Gain detects LLM-generated plagiarism via probabilistic compression, distinguishing source reuse from similarity.
Quantization precision (FP16/INT8/INT4) and prompt design critically affect biomedical LLM classifier calibration on Mistral variants.
MAFIA: query-only memory poisoning attacks on audited LLM agents via probing and factual injection.
Layer-wise analysis shows sensitivity, causality, and repair capacity dissociate when LLMs fail on perturbed input; identifies spike-and-suppress vs. late-accumulation regimes.
LatentGuard: safeguard framework compressing textual reasoning into latent states for efficient, inspectable LLM content moderation.
Computing actual causes for neural network predictions using Halpern-Pearl causal models and Boolean SCMs to improve explainability under structured input dependencies.
Faithfulness-safety tension in Large Reasoning Models: models must be interpretable for monitoring yet robust against unsafe reasoning paths.
Multi-agent clinical committees using Gemini show vulnerability to social shortcut cascades where peer consensus propagates errors across agents.
MissClick: adversarial attack on GUI grounding models exploiting digit-serialized coordinate generation to induce large spatial click displacement.
CARE-Bench evaluates 11 LLMs on medical triage safety, testing when patient-facing models recommend escalation to clinicians.
Task system for detecting hallucinations in Arabic Islamic QA and selecting verified answers from candidates.
AntiSkillBench: end-to-end benchmark evaluating privacy leakage, impersonation risk, and defenses in persona-skill pipelines for agents.
VetScore evaluates veterinary QA outputs by assessing citation faithfulness and weighting claims by medical harm potential.
DiagLoop generates counterfactual training data from clinical guidelines to improve diagnostic LLM reasoning in low-data settings.
Higher-order safety shields enforce multi-constraint safety (speed, force, jerk limits) for cyber-physical systems beyond binary state safety.
RAPO framework addresses catastrophic forgetting in continual multimodal LLM post-training via explicit dual-channel risk governance.
Study identifies spurious language priors in LLM token-level supervision during on-policy distillation and proposes mitigation.
ConformalShift demonstrates event-reordering attacks on adaptive ECG monitoring via conformal prediction without waveform modification.
Systematic study reveals gender bias in LLM-based fake news detection across six state-of-the-art models using LIAR benchmark.
Proposes security-oriented lifecycle model for LLM systems addressing provenance, signing, permissions, and decommissioning in critical infrastructure.
Formal verification framework for LLM-agent systems operating on persistent relational data with business logic constraints.
OpenAI disrupted Cambodia-based scam ring leveraging ChatGPT for investment, romance, and impersonation fraud.
Simon Willison argues against 'meat proxy' behavior—blindly relaying AI output without validation—and advocates for critical engagement and reformulation of AI-generated content.
Formalizes the missing-target problem in fairness audits: how to justify demographic distributions in open-ended generation when ground truth is undefined.
MedPRESS: 600-dialogue benchmark measuring patient-pressure-induced sycophancy in LLMs across medication, self-care, and triage scenarios.
Magnet framework detects cross-session AI misuse where attackers decompose harmful goals into innocuous agentic tasks across isolated sessions.
Study proposes longitudinal measurement framework to detect cognitive, developmental, and socio-affective behavioral changes from long-term LLM interaction rather than short-term evaluations.
Private Bayesian bootstrap technique using group-level blocking weights to privatize both point estimates and uncertainty quantification for statistical reporting.
Digital Twin-Enhanced Multiscale Planning automates incident response via decision-theoretic agents, bridging abstract models to operational systems.
Import AI newsletter curates self-sustaining AI viruses concept, AI progress pacing debate, and creativity attribution confusion.
DeBERTa-Sentinel uses disentangled attention for robust AI-generated text detection resistant to paraphrasing and model-diversity attacks.
Decoy images amplify caption-mediated defenses (ECSO) against encoded jailbreaks on VLMs, reducing attack success up to 73 percentage points.
VLAGuard framework defends VLA robots against physical adversarial patches via attention-protective fine-tuning.
Caliber defends against model extraction by adding calibrated Gaussian noise to logits with provable query costs.
LLMs exhibit medical sycophancy—abandoning correct answers under user pushback—driven by conversational factors, not fixed model properties.
Attribute-level unlearning for MLLMs enables fine-grained removal of sensitive identity information while preserving model utility.
Two-sided audit framework for self-improving AI-for-science systems to distinguish real gains from search artifacts and oracle drift.
Zero-query jailbreaks for text-to-image systems via filter-generator discrepancy; transfer-based attack requiring no target queries.
Neuro-symbolic governance framework for verifiable AI agents in decentralized digital twin ecosystems with semantic profile layers.
Calibration breaks under unseen subtype shift: models maintain accuracy but become systematically overconfident on novel fine-grained categories.
Greg Brockman (OpenAI) notes ChatGPT-Slack integration friction: users reject AI-initiated requests, prefer AI enhances rather than replaces human interaction.