nvidia-cosmos/cosmos-predict2.5 — Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models family – Code: Apache 2.0

TencentCloud/CubeSandbox — Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents. — Apache 2.0

CLARIN Café JOHD - Reviving Legacy WordNet-like Resources

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM

Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking

StayStill: a large-scale 3D idle animation dataset, project, dataset, code

poloclub/transformer-explainer — Transformer Explained Visually: Learn How LLM Transformer Models Work with Interactive Visualization

Exploring Resolution-Wise Shared Attention in Hybrid Mamba-U-Nets for Improved Cross-Corpus Speech Enhancement, code

Inverse-Hessian Regularization for Continual Learning in ASR

Bridging the Modality Gap: Softly Discretizing Audio Representation for LLM-based Automatic Speech Recognition

frankmorgner/vsmartcard — umbrella project for emulation of smart card readers or smart cards

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching, code

Modern C++ Is A Lie: Chromium Treats Half The Standard Library As A Bug

unsloth/FLUX.2-klein-4B-GGUF

Nord-Parl-TTS: Finnish and Swedish TTS Dataset from Parliament Speech

Replacing Protobuf with Rust to go 5 times faster

noctalia-dev/noctalia — A sleek and minimal desktop shell thoughtfully crafted for Wayland.

AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching, code

Whisper-MLA: Reducing GPU Memory Consumption of ASR Models based on MHA2MLA Conversion, code

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

Why Language Models Hallucinate

PRiSM: Benchmarking Phone Realization in Speech Models

WAXAL: A Large-Scale Multilingual African Language Speech Corpus, google/WaxalNLP

Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces, code

Archaeologists Unearth “First Direct Evidence” of Advanced Ancient Metallurgy in Egypt’s Middle Kingdom

Do Neural Codecs Generalize? A Controlled Study Across Unseen Languages and Non-Speech Tasks

Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data

Man Plays The Entire REIGN IN BLOOD Album At 200% SPEED In 1 Take

gaHealth: An English–Irish Bilingual Corpus of Health Data

@inproceedings{lankford-etal-2022-gahealth,
    title = "ga{H}ealth: An {E}nglish{--}{I}rish Bilingual Corpus of Health Data",
    author = "Lankford, S{\'e}amus  and
      Afli, Haithem  and
      N{\'i} Loinsigh, {\'O}rla  and
      Way, Andy",
    editor = "Calzolari, Nicoletta  and
      B{\'e}chet, Fr{\'e}d{\'e}ric  and
      Blache, Philippe  and
      Choukri, Khalid  and
      Cieri, Christopher  and
      Declerck, Thierry  and
      Goggi, Sara  and
      Isahara, Hitoshi  and
      Maegaard, Bente  and
      Mariani, Joseph  and
      Mazo, H{\'e}l{\`e}ne  and
      Odijk, Jan  and
      Piperidis, Stelios",
    booktitle = "Proceedings of the Thirteenth Language Resources and Evaluation Conference",
    month = jun,
    year = "2022",
    address = "Marseille, France",
    publisher = "European Language Resources Association",
    url = "https://aclanthology.org/2022.lrec-1.727/",
    pages = "6753--6758",
}

UCCIX: Irish-eXcellence Large Language Model

Acoustic Phonetics

Parachute use to prevent death and major trauma when jumping from aircraft: randomized controlled trial