VNIbCReg: VICReg with Neighboring-Invariance and better-Covariance Evaluated on Non-stationary Seismic Signal Time Series

Vosk models

Ten tips for writing a brilliant PhD thesis – and enjoying the process

  1. Arrange a comfortable, quiet writing space
  2. Study submission requirements
  3. Detailed table of contents
  4. Write the first sentence
  5. Make the introduction tantalising and the conclusion decisive
  6. Keep the literature review focussed
  7. Develop a disciplined scientific writing style
    • Avoid jargon
    • Avoid superlatives
    • Avoid passive voice
    • Avoid first person
    • Short, sharp sentences
    • Prefer “of” to genitive
    • Take care with verb agreement
  8. Be fastidious in your referencing
  9. Craft a succinct and informative title
  10. Ensure your abstract captures the ‘big picture’

submission site for IEEE Signal Processing Letters

Information for Authors-SPL

rhasspy/sv_kaldi-rhasspy

Do We Know Which Pre-trained Model Outperforms TIMIT Phoneme Recognition?

ML4ITS/vibcreg

BalajiAI/VICReg — JAX/Flax implementation of “VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning”. ICLR 2022.

VIbCReg: Variance-invariance-better-covariance regularization for self-supervised learning on time series, Computer Vision Self-supervised Learning Methods on Time Series (updated title)

kaldi_swe, huggingface

Fast vs Slow Swedish


Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning, code (Embedding model is open, “Perception LM” is not), Perception Encoder Audio-Visual (PE-AV)

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

ML4ITS/repositories

Reinventing Entropy - Compression is Intelligence Part 1

DSTA: Reinforcing Vision-Language Understanding for Scene-Text VQA With Dual-Stream Training Approach

GigaAM Multilingual & GigaChat Audio — accepted to InterSpeech 2026, released fully open:

  • ai-sage/GigaAM-Multilingual
    • MIT
    • Conformer-based foundation models (220M / 600M parameters)
    • HuBERT-style pre-training objective
    • trained on 2M hours of speech across 70+ languages
    • fine-tuned character CTC decoders on 50K hours
    • Fine-tuning guide
  • GigaChat Audio 10B
    • MIT
    • built on GigaChat
    • Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder
  • ai-sage/TimeGround-1M
    • CC-BY
    • Synthetic English audio dataset for time-aware speech understanding
    • temporal localization, temporal description, and timed summaries
    • Based on YODAS2 English

Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

wavtechyukky/pyshiro

Harnessing Whisper for Prosodic Stress Analysis

@inproceedings{sohn-etal-2025-harnessing,
    title = "Harnessing Whisper for Prosodic Stress Analysis",
    author = "Sohn, Samuel S.  and
      Knutsen, Sten  and
      Stromswold, Karin",
    editor = "Che, Wanxiang  and
      Nabende, Joyce  and
      Shutova, Ekaterina  and
      Pilehvar, Mohammad Taher",
    booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
    month = jul,
    year = "2025",
    address = "Vienna, Austria",
    publisher = "Association for Computational Linguistics",
    url = "https://aclanthology.org/2025.findings-acl.1331/",
    doi = "10.18653/v1/2025.findings-acl.1331",
    pages = "25931--25942",
    ISBN = "979-8-89176-256-5",
}

Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis

ZIM file format

MFA 2.2.15 Swedish Waxholm models

Inverse-Hessian Regularization for Continual Learning in ASR, arXiv

@INPROCEEDINGS{11461503,
  author={Vander Eeckt, Steven and Van Hamme, Hugo},
  booktitle={ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)}, 
  title={Inverse-Hessian Regularization for Continual Learning in ASR}, 
  year={2026},
  volume={},
  number={},
  pages={18192-18196},
  doi={10.1109/ICASSP55912.2026.11461503}
}