Interesting links, 20/07/2026
Misc. interesting things.
Ten tips for writing a brilliant PhD thesis – and enjoying the process
- Arrange a comfortable, quiet writing space
- Study submission requirements
- Detailed table of contents
- Write the first sentence
- Make the introduction tantalising and the conclusion decisive
- Keep the literature review focussed
- Develop a disciplined scientific writing style
- Avoid jargon
- Avoid superlatives
- Avoid passive voice
- Avoid first person
- Short, sharp sentences
- Prefer “of” to genitive
- Take care with verb agreement
- Be fastidious in your referencing
- Craft a succinct and informative title
- Ensure your abstract captures the ‘big picture’
submission site for IEEE Signal Processing Letters
Do We Know Which Pre-trained Model Outperforms TIMIT Phoneme Recognition?
BalajiAI/VICReg — JAX/Flax implementation of “VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning”. ICLR 2022.
VIbCReg: Variance-invariance-better-covariance regularization for self-supervised learning on time series, Computer Vision Self-supervised Learning Methods on Time Series (updated title)
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning, code (Embedding model is open, “Perception LM” is not), Perception Encoder Audio-Visual (PE-AV)
PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reinventing Entropy - Compression is Intelligence Part 1
GigaAM Multilingual & GigaChat Audio — accepted to InterSpeech 2026, released fully open:
-
ai-sage/GigaAM-Multilingual
- MIT
- Conformer-based foundation models (220M / 600M parameters)
- HuBERT-style pre-training objective
- trained on 2M hours of speech across 70+ languages
- fine-tuned character CTC decoders on 50K hours
- Fine-tuning guide
-
GigaChat Audio 10B
- MIT
- built on GigaChat
- Conformer speech encoder and a modality adapter feed audio embeddings directly into a Mixture-of-Experts decoder
-
ai-sage/TimeGround-1M
- CC-BY
- Synthetic English audio dataset for time-aware speech understanding
- temporal localization, temporal description, and timed summaries
- Based on YODAS2 English
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
There was a 'Moved Permanently' error fetching URL: 'https://x.com/4wavetech/status/2079258991528726586'
Harnessing Whisper for Prosodic Stress Analysis
@inproceedings{sohn-etal-2025-harnessing,
title = "Harnessing Whisper for Prosodic Stress Analysis",
author = "Sohn, Samuel S. and
Knutsen, Sten and
Stromswold, Karin",
editor = "Che, Wanxiang and
Nabende, Joyce and
Shutova, Ekaterina and
Pilehvar, Mohammad Taher",
booktitle = "Findings of the Association for Computational Linguistics: ACL 2025",
month = jul,
year = "2025",
address = "Vienna, Austria",
publisher = "Association for Computational Linguistics",
url = "https://aclanthology.org/2025.findings-acl.1331/",
doi = "10.18653/v1/2025.findings-acl.1331",
pages = "25931--25942",
ISBN = "979-8-89176-256-5",
}
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
MFA 2.2.15 Swedish Waxholm models
Inverse-Hessian Regularization for Continual Learning in ASR, arXiv
@INPROCEEDINGS{11461503,
author={Vander Eeckt, Steven and Van Hamme, Hugo},
booktitle={ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
title={Inverse-Hessian Regularization for Continual Learning in ASR},
year={2026},
volume={},
number={},
pages={18192-18196},
doi={10.1109/ICASSP55912.2026.11461503}
}