Machine Translation between English and Syriac (East Syriac Dialect) using Statistical Machine Learning
First phrase-based statistical MT system for Syriac (East dialect) addresses UNESCO endangered-language gap via novel corpus.
Search the full wire by company, model, lab, or keyword. Every story we have ever aggregated.
First phrase-based statistical MT system for Syriac (East dialect) addresses UNESCO endangered-language gap via novel corpus.
Spectral method with provable guarantees for node label recovery on noisy signed graphs via principal eigenvector decoding.
TriProbe framework diagnoses task separability across input, feature, and classifier levels via multi-level probing for model interpretability.
VoiceTrace benchmark enables speaker-aware speech retrieval combining semantic content search with speaker identification from reference utterances.
AeroWeaver integrates LLM agents into UAV swarms for high-level planning and distributed autonomous coordination across multi-agent systems.
COMPASS-ABS reduces GPU cluster fragmentation for deep learning training via adaptive scheduling maintaining low fragmentation without workload priors.
Aligned Continuous Integrate-and-Fire framework compresses speech tokens efficiently for zero-shot SpeechLLM alignment without full-model fine-tuning.
Cunning questions training enhances LLM safety vigilance by teaching detection of latent risks hidden in benign contexts beyond surface compliance.
ActiveScale framework enables vision-language-action models to perform active perception via coordinated model, data, and hardware designs for robotic manipulation.
LDE framework leverages collaborative perception data as supervision for unsupervised domain adaptation in autonomous driving perception models.
MiST suite (8B/32B models) applies mid-training on curated cybersecurity corpus for strong domain performance without raw continual pre-training.
154M-parameter HTML-aware foundation model for Czech web document representation with ModernBERT architecture.
Analysis of kinetic energy spectra in ML weather models (NeuralGCM, FourCastNet, AIFS, GenCast) vs. physics-based IFS.
Contrastive Noise Alignment improves diffusion/flow-matching training by dynamically optimizing data-noise couplings.
ActionPiece rethinks action tokenization in vision-language-action models to preserve task-relevant adjustments beyond MSE.
Hyperbolic embeddings applied to biomedical knowledge graphs for Mendelian disease differential diagnosis.
TypeSafe releases Jev, a lightweight classification/routing model claiming 100x+ speed and 200x+ cost reduction vs. frontier LLMs.
Large Reasoning Models exhibit Onset Refusal Collapse at token onset; token-level analysis reveals safety alignment vulnerabilities.
Mixture-of-Bottleneck Experts reformulates video sentiment analysis as ordinal regression across text, audio, image modalities.
Spatially adaptive noise injection in diffusion samplers applies stochastic correction selectively based on image geometry.
Causal Semantic World Action Model augments FastWAM with V-JEPA 2.1 for robust visual OOD generalization in action inference.
LGM applies neuro-symbolic reasoning to disentangle long-term memory for context-dependent personalized agent reasoning.
Multi-agent LLM systems can propagate unsafe behaviors through communication; rare local deviations scale to collective failure when contagion outpaces correction.
Vision-Language Models report high confidence despite incorrect reasoning; verbalized confidence is trajectory-independent and insensitive to reasoning quality.
Agent skill libraries are English-dominant; M-SQE synthesizes multilingual skills to improve retrieval accuracy for low-resource languages like Swahili and Hindi.
Amazon is letting all customers use Alexa+ assistant in early access period
RiskWorld framework fuses spatial risk fields with occupancy forecasting for safer automated driving trajectory planning and selective replanning.
Delphos applies multitask RL to learn transferable discrete choice model specification strategies across transport datasets.
Tests whether Claude 3.5 Haiku's rhyme planning via newline-resident features generalizes across seven open-weights models and six cross-layer transcoders.
WetRobo is a reproducible robot kit for wet-lab automation that transfers vision-language-action policies between laboratory environments without teleoperation.