TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling
TransBERT framework pre-trains domain-specific NLP models using synthetic translations; releases TransCorpus toolkit for French life sciences.
Every story tagged with this topic, ordered by date.
TransBERT framework pre-trains domain-specific NLP models using synthetic translations; releases TransCorpus toolkit for French life sciences.
Xiaomi releases MiMo-V2.6-Pro 1T-A42B open-weights model trained for $3M, claiming top performance among open models.
Fine-tuning Qwen2.5-7B and Ministral-8B on personality-labeled corpora improves consistency vs. instruction prompting for social agents.
Open-weight LLMs tested on Norwegian Moral Foundations Questionnaire; persona steering and activation-level intervention techniques partially shift moral profiles.
dQwen3.5 adapts Qwen 0.8B–9B hybrid-attention models to diffusion language models by bidirectionalizing RNN layers without full retraining.
154M-parameter HTML-aware foundation model for Czech web document representation with ModernBERT architecture.
Tests whether Claude 3.5 Haiku's rhyme planning via newline-resident features generalizes across seven open-weights models and six cross-layer transcoders.
JustFit enables 200K-token LLM inference on 24GB laptop via just-in-time state management and KV compression.
MéTRON-FR: 125M GPT-2 model pretrained on French scores 85.97% on QFrBLiMP; reveals tokenizer sensitivity in cross-lingual evaluation.
OPEN-1B: fully auditable 1B-parameter LLM training run with hardware-agnostic reproducibility verification against undisclosed data/backdoors.
Nameless tokenization defends open-weight LLMs against control-token forgery, fixing vulnerability in all 256 audited chat tokenizers.
Cohere releases North Small Translate, an open-weight machine translation model optimized for speed and cost efficiency.
Rosetta system fine-tunes NileChat-3B with LoRA for English-to-Dialectal Arabic dialogue translation, achieving 4th place in AlexandriaX-2026 shared task.
AuK, open-source multimodal model for unified speech generation and editing via natural-language instructions and audio context.
Study characterizes interaction between inference-time activation steering and weight-only quantization (INT8, NF4) on 7-9B open-weight LLMs.
Study of 3,471 uncensored open-weight models on HuggingFace (Jan 2024–Mar 2026), tracking safety guardrail removals and redistribution persistence.
Causal taxonomy distinguishes deceptive behavior from deceptive mechanisms in language models, tested on open-weight families.
Open-weight transformers show logical validity representations remain decodable from hidden states despite near-chance behavioral performance.
Google announces August 2026 AI updates; article lacks specific details on Gemma releases or capabilities.
Causal analysis across 9 open-weight models shows quantization damage is distributed globally, not localized to task circuits or weight statistics.
HiveTraceGuard-Pro, a 0.6B LoRA guardrail model tuned on Qwen3, detects Russian and English prompt injection and jailbreaks with binary safety scoring.
Tencent releases Hy4, a 770B open-weight LLM with 49B active params and 1M token context, 2.6× larger than Hy3.
Open-source tendon-driven robotic hand simulator for dexterous manipulation learning with underactuated transmission dynamics.
Three-stage post-training recipe (acquire, repair, preserve) for 2B open-weight dialogue game agents using diagnostic error analysis and RL.
Linear probe analysis of how open-weight LLMs organize Moral Foundations Theory categories in representation space.
Puro-2B: open-source 1.5B LM trained on RTX 5090 for <$5k, targeting cost-accessible pretraining.
Qwen releases Qwen3.8-Flash-Next, a 125B-parameter MoE model with 6B active tokens and multimodal capabilities, previewing Qwen4 architecture.
Z.ai CEO Jie Tang discusses GLM 5.3 model and post-training scaling laws as alternative to parameter growth.
Glean CEO explains model routing as cost-control mechanism driven by frontier model pricing and open-weights adoption, improved via scaled human feedback.
Mojo programming language released as open source under Apache 2 license after 1.0 stable release, pivoting from Python superset goal.
Google announces July 2026 AI updates; article lacks specific details on models, features, or benchmarks.
Alibaba releases Qwen 3.8 Max (2.4T params) and 27B open-weight models optimized for coding and collaboration tasks.
Analysis of 18 open-source LLMs showing cultural bias in mythology knowledge; models encode cross-cultural distinctions in residual streams but fail to decode non-Western traditions.
Antares: compact LLMs (350M–3B) for agentic vulnerability localization via SFT and RL on cybersecurity reasoning over code.
Microsoft-led open letter signed by 235 AI companies including NVIDIA and OpenAI argues against US government restrictions on open-weight models on safety grounds.
Gaokerena: compact Persian-language medical LLM family trained on 90M-token corpus for low-resource healthcare deployment.
Tevatron 3.0 integrates Megatron-Core for efficient MoE reranker training, enabling billion-scale cross-encoder + distillation workflows on academic budgets.
DeepSeek releases V4-Flash-0731, a 304B parameter model with enhanced agentic capabilities, outperforming larger competitors at $0.14/$0.27 per million tokens.
Podcast discussion on open-weight model competitiveness, cybersecurity risks, and AI leadership policy letters signed by industry leaders.
Simon Willison releases smevals, an open eval framework for benchmarking models, prompts, and inference harnesses across configurations.
DenseOn and LateOn: open-source 149M-parameter retrieval models trained on 1.88M supervised pairs; competitive on multilingual and code search.
Method to detect CSAM-generating LoRAs from weight fingerprints (singular vectors) without generating outputs, enabling safer moderation.
Kimi K3 open-weights model released amid broader industry discussion on open model availability and strategy.
Moonshot releases Kimi K3 weights (2.8T params, 1.56TB) with modified MIT license requiring attribution for products >100M MAU.
Anthropic publishes official stance on open-weights model releases, addressing trade-offs between transparency, safety, and competitive positioning.
Causal-TS: open-source Python library for causal discovery in high-dimensional nonstationary time series with GPU-accelerated conditional independence testing.
Kimi K3: 2.8T MoE model with 104B active params, 1M context window, Delta Attention, 2.5x scaling efficiency over K2.
ELMOD: 2.7B German-language model optimized for mobile deployment using public data and morphology-aware preprocessing.
Byte-Prefix Marginalization method for cross-tokenizer on-policy distillation of open-weight LLMs with incompatible vocabularies.
OpenForgeRL enables end-to-end training of harness-native agents with open infrastructure, addressing limitation of complex inference harnesses like Claude Code.