ChatGPT starts blocking direct requests to copy an author's style
New behavior capturing a writer's "broad qualities" could have legal implications.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
New behavior capturing a writer's "broad qualities" could have legal implications.
People visit the booth of Kimi, an LLM developed by the Chinese startup Moonshot, during the World AI Conference in Shanghai, China, July 20th. | Image: LONG WEI/ Feature China/Future Publishing via Getty Images Silicon Valley has spent much of the past week on red alert, digesting the arrival of Moonshot AI's Kimi K3, a Chinese AI model that can allegedly beat some of the best systems built by US companies at a fraction of the cost. Its performance alone would have been enough to intensify the rivalry between the US and China. But Moonshot's plan to release the model's weights for free - and...
Kimi K3: 2.8T MoE model with 104B active params, 1M context window, Delta Attention, 2.5x scaling efficiency over K2.
Study on attribution hallucination in VLMs: models answer correctly but fail to identify supporting evidence regions without coordinate supervision.
Framework for auditing LLM social simulators by validating reasoning patterns, not just output match, via sunscreen concept test.
Argument that autonomous research systems need efficiency metrics alongside quality; search cost becomes critical as AR scales to expensive domains.
Meta on Monday said it is rolling out its Meta AI chatbot within Threads' DMs, giving users a way to chat with the AI assistant.
Study of sparse autoencoders' downstream causal effects via logit geometry; shows activation descriptions don't predict steering behavior.
APPA: IFC framework for LLM agents handling mixed-confidentiality data via context branching and prospective access control against injection attacks.
Crowdsourcing label aggregation model for imbalanced classes with class-dependent annotator accuracy; addresses minority-class detection in inspection systems.
1D CNN for gyroscope bias correction in spacecraft with aleatoric and epistemic uncertainty quantification via ensemble.
Study of correctness degradation in code repair agents under forced revision loops; 30 HumanEval benchmarks show revision doesn't guarantee reliability.
User study (n=34) on explainable AI impact for LLM code review: XAI effects on developer trust in automated review reasoning.
NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they... NVIDIA Ising Calibration is an open source vision language model (VLM) designed to interpret diagnostic outputs from quantum processors and determine how they should be tuned to continue operating. This post introduces the latest model release, NVIDIA Ising Calibration 1.5, which advances AI-based QPU calibration by analyzing unfamiliar diagnostic results without prior training examples. Source
PIVOT reduces token-level sparse attention indexing cost via query-group caching, addressing bottleneck in DeepSeek's production sparse attention.
Google’s AI Overviews now appear in 43% of searches, underscoring how quickly AI-generated answers are becoming the default way people discover information online.
Survey of AI's role in innovation ecosystems covers digitization trends and macro economic integration.
SIREN applies LLM agents to end-to-end extreme-weather early warning with multi-step decision workflows.
D-Score detects hallucinations via spectral analysis of hidden activations, using single forward pass.
ELMOD: 2.7B German-language model optimized for mobile deployment using public data and morphology-aware preprocessing.
PYPM-GGD extends Pitman-Yor nonparametric Bayesian methods to non-conjugate posteriors via variational inference.
CADER adaptively allocates vision-language model compute for long-video QA based on per-example confidence.
Evaluates fuzz testing methods for safety-critical RL agents across robotics and autonomous systems with standardized metrics.
LLM-SoccerArena: live prospective benchmark measuring LLM forecasting on real-world sports outcomes pre-event.
Adapts MLLMs (Qwen3-VL) to sparse 8-16 frame video input for temporal grounding, addressing training-deployment mismatch.
Reduced-order models for wake flow control trade compactness against forecast accuracy using latent-space encoders.
FPGA evaluation of learned feature gating in fixed-point automatic modulation classification with quantization trade-offs.
Deep semantic hashing method DSCH-Loss for efficient approximate nearest neighbor search in high-dimensional spaces.
TRACE-CTI framework audits cyber threat intelligence mappings to MITRE ATT&CK with provenance tracking and consensus validation.
Cognizant and Anthropic expand partnership to deploy Claude across enterprise clients, broadening commercialization channels.