Open tabs, 20/04/2026
Misc. interesting things.
- Irish Language & Texts
- Phonetics & Acoustic Analysis Tools
- ASR / Speech Recognition
- Speech Synthesis & TTS
- G2P & Phoneme Processing
- Embedding Models & Sentence Transformers
- NLP / LLM
- Swedish, Sami & Nordic Languages
- Hungarian Language
- Conferences & Journals
- Personal Work
- Parliamentary & Text Corpora
- LibriVox & Project Gutenberg
- Entertainment & Media
- Maths & Foundations
- Misc
Irish Language & Texts
Phonological Resources
- Irish Pronunciation Database: sroich
- ceannaigh - Wiktionary
- sroich - Wiktionary
- croith - Wiktionary
- muintir - Wiktionary
- Cross-dialect lexicon optimisation for an endangered language ASR system: the case of Irish
- ifaa2020.pdf (Kevin Scannell)
- Object detection for cross-linguistic vowel analysis
Wikisource
- A Dialect of Donegal/Texts/Leadairt na bhfear mór
- Index:Die araner mundart.djvu
- Seite:Die araner mundart.djvu/26
- Seite:Die araner mundart.djvu/28
- Page: A contribution to the phonology of Desi-Irish …/13
- Index:A contribution to the phonology of Desi-Irish
- Leabhar Sgeulaigheachta/Monachar agus Manachar
Leigh Leat
- Turas Róise - Leathanach 4
- AN GHAOTH AGUS AN GHRIAN
- Leabhairiní - An Trá
- Clíona agus an Claíomh - Leathanach 2
- Lá Breithe Sona Dhuit!
- A BHUÍ! Cá bhfuil tú, a Bhuí?
- Abair Liom, An Rud Maith é an Drón?
phonlab-tcd
- phonlab-tcd repositories
- phonlab-tcd/Radio-Archivist
- phonlab-tcd/backoffice
- phonlab-tcd/Qomhra/Corpus_V2
- phonlab-tcd/litreoir: Irish spelling test app
- phonlab-tcd/mile-glor-na-nog
- phonlab-tcd/ASR-Whisper-Large
- MengjieQian/fairseq
- NeasaNi/bat-mirialta
- JohnSloan8/abair
Other Irish Resources
- Corpas na Gaeilge
- Scéalta & Leabhair (Cloich Cheann Fhaola)
- Rang 5-6 - Robo
- Leabhair do Pháistí (padlet)
- Teanga Tí - Glór na nGael
- Fís & Fuaim - Glór na nGael
- An tIriseoir by Michelle Nic Pháidín - Goodreads
- ClubLeabhar: I dtír mhilis na mbeo
- ClubLeabhar: June 2017 extract
- ClubLeabhar: Idir Dhá Thír
- Géibheann - reading (YouTube)
- Géibheann (YouTube)
Other Irish Tools
Phonetics & Acoustic Analysis Tools
WaveSurfer & Snack (KTH)
- WaveSurfer - Wikipedia
- WaveSurfer.js - audio waveform player
- Snack Home Page
- kth-tmh/snack
- snack/tests/all.tcl
- snack/python
- snack/python/snack
- PR #4: Add GitHub Actions build/release workflow
- PR #6: Debian
- PR #8: Python ext
- PR #9: Merge main into python-ext
- PR #12: NumPy/librosa/torchaudio interoperability
- Debian – wavesurfer in sid
- GUB Facial Animation Player
- phonerec/phonerec.py
- tcl-snack patch for Windows
Formant Analysis
- tabahi/formantfeatures
- GuiMarion/formantsExtraction
- How to extract formant tracks with Praat and Python
- Comparative Evaluation of Acoustic Feature Extraction Tools for Clinical Speech Analysis
- Evaluating the Representation of Vowels in Wav2Vec Feature Extractor: A Layer-Wise Analysis Using MFCCs
- Optimizing the Extraction of Vowel Formants (ResearchGate)
- A Robust Formant Extraction Algorithm Combining Spectral Peak Picking and Root Polishing
- p13.3_130.pdf (ICPhS 1995)
- Design, analysis and experimental evaluation of block based transformation in MFCC computation for speaker recognition
- Acoustic Theory of Vowel Production - Ento Key
Other Phonetics Tools
- pyAudioAnalysis
- CPJKU/madmom
- libAudioFlux/audioFlux
- Nitnelav/awesome-acoustic
- timmahrt/pyAcoustics/nucleus_detection_matlab
- MitchellAcoustics/Soundscapy
- pyfar.signals documentation
- AcousTools: Python-Based Acoustic Holography Library
- phonlab/phonlab/utils/tidy.py
- Penn Phonetics Laboratory
- pysptk introduction (nbviewer)
- pysptk Speech analysis and re-synthesis (nbviewer)
- Speech-Articulatory-Coding/sparc/sparc.py
- VocalTractLab
- Multilingual speech-to-vocal tract visualization (paper)
- MultiSpeechToVocalTract/code/create_video.py
- Siteseeker Voice (Swedish speech search)
KTH Speech History / Gunnar Fant
- KTH STL-QPSR Archive
- Gunnar Fant > Vocal Pedagogy
- The OVE cascade formant synthesizer of Gunnar Fant, 1953
- The OVE III speech synthesizer (PDF)
- De som lyssnar till maskiner (PDF)
- Klatt review of TTS for English (web archive)
- Halfacentury.pdf (Fant)
- gunnar_is2009.pdf
- Speech transmission device - Tekniska museet
- QPSR 1960 issues
- QPSR 1960: Detection of voicing and automatic segmentation schemes
- QPSR 1960: Formant frequency tracking
- QPSR 1960: Pole-zero matching techniques
- QPSR 1960: RASSLAN — a 6-channel closed loop sectioning device
- QPSR 1960: Formant-tracking
- QPSR 1960: Structural classification of Swedish phonemes
- QPSR 1965: A filter bank speech spectrum analyzer
- joregan/QPSR: 1969/10/1 PDF
- Basic Equipment (Springer — Fant chapter)
- Speech Acoustics and Phonetics: Selected Writings (review, Project MUSE)
ASR / Speech Recognition
CTC Decoding
- ASR Inference with CTC Decoder — Torchaudio 0.12.0
- ASR Inference with CTC Decoder — Torchaudio 2.8.0
- CTCDecoder — Torchaudio 2.9.0
- torchaudio.models.decoder._ctc_decoder
- asr_inference_with_ctc_decoder_tutorial.ipynb (Colab)
- Torchaudio Documentation — 2.9.0
Models & Frameworks
- facebookresearch/omnilingual-asr
- omnilingual-asr/wav2vec2_llama/syntax.py
- omnilingual-asr/norm_config_module.py
- omnilingual-asr/wav2vec2/asr/README.md
- Omnilingual ASR: Open-Source Multilingual Speech Recognition (HuggingFace collection)
- NeMo/conformer_transducer_char.yaml
- NeMo/k2/loss_mixins.py
- Entropy-Based Methods for Word-Level ASR Confidence Estimation (NVIDIA)
- waveletdeboshir/gigaam-ctc
- salute-developers/GigaAM
- GigaAM transformers gigaam_transformers.py
- lumaku/ctc-segmentation
- MFA multiprocessing.py
- scb10x/monsoon-whisper-medium-gigaspeech2
- rVAD/rVADfast_py_2.0
- ScholarOne Manuscripts (IEEE SPS)
- How to Generate Text in One Step
- ISCA Archive: Podcastle (Interspeech 2009)
Speech Corpora
- Common Voice Scripted Speech 25.0 - English
- Common Voice Scripted Speech 25.0 - Romansh Vallader
- common-voice/cv-dataset
- What happened to Common Voice on HuggingFace? (Reddit)
- fsicoli/common_voice_15_0 ga-IE test.tsv
- Kathleen 1.0 - Mozilla Data Collective
- Joe 1.0 - Mozilla Data Collective
- Mozilla Data Collective datasets (English)
- KTH/speechdat
- KTH/nst
- hifitts-2
- PleIAs/YouTube-Commons
- NeMo speech data processor
- CORAAL config.yaml (NeMo)
- PleIAs/YouTube-Commons README
- NeMo-speech-data-processor: Granary PR #135
- yfyeung/vctk
- vctk - TensorFlow Datasets
- TSP Speech Database
- TSP Lab - Data
- Standardised reading (York)
- openslr/openslr dataset
- openslr.org/12 (LibriSpeech)
- openslr.org/152
- kaggle/libritts dataset
- NaoyukiKanda/LibriSpeechMix
- BEAT dataset
- H-Liu1997/BEAT (HuggingFace)
- datashare.ed.ac.uk/handle/10283/2120 (CSTR)
- EARS Dataset
- Grosy/wav2vec2-base-hu
- ctaguchi/ikema_youtube_asr_full dataset
- ctaguchi (Chihiro Taguchi) models
- HPLT datasets v3.0, Hugging face
- NetherlandsForensicInstitute/stackexchange-duplicate-questions-translated-nl
- SLAM-Omni: worstchan collection
- worstchan/VoiceAssistant-400K-SLAM-Omni
- MrDragonFox/Elise dataset
- MediaEval 2013 SWS submission (CEUR)
- MediaEval 2014 submission (CEUR)
- MultimediaEval datasets
- Scotland 5 - IDEA: International Dialects of English Archive
- IDEA Copyright & Credit Information
- Macquarie Dictionary Pronunciation Key
Speech Synthesis & TTS
TTS Models
- ResembleAI/Chatterbox app.py
- OpenBMB/VoxCPM: VoxCPM2 tokenizer-free TTS
- VoxCPM2 on ModelScope
- kyutai-labs/pocket-tts
- kyutai/pocket-tts
- kyutai/tts-voices
- neuphonic/neutts-air
- microsoft/VibeVoice-ASR
- Orpheus TTS notebook
- nv-tlabs/kimodo
- nvidia/Kimodo space
- Unsloth TTS fine-tuning guide
- unslothai/unsloth
- Wataru-Nakata/miipher
- Shinnosuke Takamichi - jvs_corpus
- DSA-Tokenizer: Disentangled Semantic-Acoustic Tokenization via Flow Matching-based Hierarchical Fusion
- Frame-Wise Breath Detection with Self-Training: An Exploration of Enhancing Breath Naturalness in Text-to-Speech
- Respiro-en/demo.ipynb
- Scaling Open Discrete Audio Foundation Models with Interleaved Semantic, Acoustic, and Text Tokens
- marin/experiments/audio
- potsawee/marin
- LaFresCat: A studio-quality Catalan multi-accent speech dataset for text-to-speech synthesis
- TGuichoux/Gelina
- Intrasentential English in Swedish TTS (KTH DiVA)
- Intrasentential English in Swedish TTS: perceived English-accentedness
Voice Conversion / Disentanglement
- auspicious3000/SpeechSplit
- Unsupervised Speech Decomposition via Triple Information Bottleneck
- auspicious3000/contentvec
- auspicious3000/AutoPST
- ContentVec paper (ICML 2022)
- SpeechTripleNet: End-to-End Disentangled Speech Representation
- An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep Learning
- SoundChoice: Grapheme-to-Phoneme Models with Semantic Disambiguation
- Towards Disentangled Speech Representations
- MixSpeech: Cross-Modality Self-Learning with Audio-Visual Stream Mixup for Visual Speech Translation and Recognition
- THE EDINBURGH INTERNATIONAL ACCENTS OF ENGLISH CORPUS: TOWARDS THE DEMOCRATIZATION OF ENGLISH ASR
- THE CMU ARCTIC SPEECH DATABASES
- Measuring the perceptual effects of modelling assumptions in speech synthesis using stimuli constructed from repeated natural speech
- Representation Learning with Contrastive Predictive Coding
- Enhancing Word Discrimination and Matching in Query-by-Example Spoken term detection with Acoustic Word Embeddings
- TranSpeech: Textless NAR Speech-to-Speech Translation
- Probabilistic Speech & Motion Synthesis (KTH DiVA)
- dspinellis/speak: Reviving the Research Edition Unix speak command
Speech Enhancement
- alibabasglab/MossFormer2
- modelscope/ClearerVoice-Studio
- AudioSet-tools: a Python framework for taxonomy-aware AudioSet curation and reproducible audio research
G2P & Phoneme Processing
- CharsiuG2P README (pretrained models)
- CharsiuG2P notebooks
- CharsiuG2P/notebooks/train_individual.ipynb
- charsiu (charsiu) models
- DeepPhonemizer/dp/model/model.py
- DeepPhonemizer Training_Example.ipynb (Colab)
- mhulden/pyfoma: Python Finite-State Toolkit
- rustfst
- Help:Pronunciation respelling key - Wikipedia
- Type IPA phonetic symbols - online keyboard
- The International Phonetic Alphabet (Cambridge)
- PRiSM: Benchmarking Phone Realization in Speech Models
- changelinglab/prism/zipa_ctc_inference.py
- Koel Labs IPA Transcription Datamodule (changelinglab/prism commit)
- opengrammar/jam-learners-grammar: Jamaican language reference
- Open Grammar Project
- Knowledge of language origin improves pronunciation accuracy of proper names
Embedding Models & Sentence Transformers
- sentence-transformers/trainer.py
- sentence-transformers/embedding-training-data
- Sentence Transformers joins Hugging Face (blog)
- SentenceTransformers Documentation
- Paper: SimCSE - Simple Contrastive Learning of Sentence Embeddings
- sentence-transformers/LaBSE
- Paper: Sentence-BERT
- How to train the best embedding model in the world
- JHU-CLSP/mmBERT
- jhu-clsp/mmBERT-base
- JHU-CLSP/ettin-encoder-vs-decoder
- Yushi-Hu/Acoustic-Span-Embeddings
- Attention-Based Audio Embeddings for Query-by-Example (Semantic Scholar)
- google-research/mseb/results.yml
- Adds a speech-to-text encoder based on LiteLLM (mseb commit)
- facebookresearch/seamless_communication
- facebookresearch/fairseq2
- SeamlessM4T-v2 (HuggingFace docs)
- alexa/massive/mt5_ic_sf_encoder_only.py
- monologg/JointBERT
- Zvec quickstart
- infiniflow/ragflow
- RAGFlow
- BAAI/bge-m3 (already covered)
- Meta AI Perception Encoder Audiovisual (PE-AV)
NLP / LLM
Language Modelling / Training
- google-research/arxiv-latex-cleaner
- google-research/speculative_kd
- LLMs-from-scratch/ch05/13_olmo3
- Reasoning Language Models: A Blueprint
- facebookresearch/matrix: Multi-Agent Data Generation
- gpt-oss-recipes/generate_flash_attention.py
- Fine-tuning GPT-OSS with Quantization-Aware Training (NVIDIA)
- Olmo 3 fully open LLM (Simon Willison)
- state-spaces/mamba
- google-research/specinvert
- google-research/dictionary_learning
- locuslab/torchdeq: Modern Fixed Point Systems
- TorchDEQ-tutorial (Colab)
- ggml-org/llama.cpp: guide running gpt-oss
- Distilling the Knowledge in a Neural Network
- HuggingFaceFW/finetranslations
- fineweb-2/fineweb2-language-distribution.csv
Models
- google/gemma-4-E4B-it
- litert-community/gemma-4-E2B-it-litert-lm
- google/tunix/models/gemma4/sampling_example.ipynb
- A Visual Guide to Gemma 4 (Maarten Grootendorst)
- LLMs-from-scratch/ch05/17_gemma4
- transformers/model_doc/gemma4.md
- ml-explore/mlx-lm
- gemma4 on ollama
- k2-fsa/OmniVoice
- OmniVoice code
- Qwen3-Next collection
- qwen3.5:9b on ollama
- google-deepmind/gemma library
- NVlabs/Jet-Nemotron
- NVIDIA Nemotron V2 collection
Tools / Infrastructure
- Onyx AI - Open Source Enterprise Search
- Onyx Documentation - Resourcing
- OpenClaw — Personal AI Assistant
- openclaw/openclaw
- openclaw/clawhub: Skill Directory
- ClawHub skills
- OpenRouter
- Arthur-Ficial/apfel: Apple Intelligence CLI with FoundationModels
- anomalyco/opencode: open source coding agent
- github/github-mcp-server
- MCP - Model Context Protocol Perl SDK
- mojolicious/mojo-mcp/t/apps/lite_app.pl
- Build an MCP server - Model Context Protocol
- Example Servers - Model Context Protocol
- servers/src/memory (MCP)
- langchain-ai/langchain
- fabric/fabric: Simple, Pythonic remote execution
Swedish, Sami & Nordic Languages
Sami Languages
- Sameradion: Stivra sivahuvvo bealálašvuođas
- Oddasat Nyheter 27 november 2025
- Oddasat Nyheter 24 november 2025
- Tablå Sameradion
- North Sami pronunciation guide (oahpamuinna blog)
- Northern Sami pronunciation dictionary (Forvo)
- Northern Sami - the days of the week (YouTube)
- Northern Sami - the weather (YouTube)
- An analysis of North Saami gradation (JSTOR via KTH)
- decentius.hit.uib.no (CG/FST test)
- duhát — Wiktionnaire (North Sami)
- čuođi — Wiktionnaire
Meänkieli / Kveeni
- Meänraatio: Valtiolta täyskäänös jahti- ja kalastusoikeuksista
- Meänraatio: Gun Olofssonin taistelusta äänen kans
Swedish
- Appendix:Swedish pronunciation - Wiktionary
- KBLab/svwikiqa dataset
- Tablå P4 Göteborg
- Riksdagstryck – Wikipedia (sv)
- swerik-project/TEI
- SweTerror: Terrorism in Swedish politics
- Terrorism in Swedish politics (SweTerror)
- Språkbanken - Språkbanken
- Ressurser fra ressursbanken (leksikon)
- asr-rapport 2024 (Norwegian)
- repo.clarino.uib.no/xmlui/handle/11509/137
- Den judiska litteraturen och översättningens dilemma (Sveriges Radio)
- Word Differences Between 4 Germanic Languages (YouTube)
- Old Norse - Can Norwegian, Danish and Swedish Speakers Understand It? (YouTube)
- Kielipankki LAT service closing (Finnish)
Hungarian Language
Learning
- Hungarian verb conjugations : r/hungarian
- books/other resources to learn Hungarian : r/hungarian
- Hungarian resources (Canva doc)
- The Top Language Resources to Learn Hungarian (Fluent Forever)
- Easy Hungarian - Patreon
- Magyar szleng : r/hungary
- Wikikönyvek:Szószedet
- tanpr57.pdf (Hungarian phonetics)
- véletlenül - Wiktionary
Mór Jókai
- Mór Jókai - Wikidata
- Mór Jókai - Wikisource
- Autor:Mór Jókai - Wikiźródła
- Timar’s Two Worlds - Project Gutenberg
- Peter the Priest - Project Gutenberg
- The Strange Story of Rab Ráby - Project Gutenberg
- The Strand Magazine: Barak Hageb and his Wives - Wikisource
- Papuga - Wikiźródła
- Magyar Elektronikus Könyvtár (libri.hu fantasy)
Conferences & Journals
Call for Papers / Out of date
- TSD 2026
- TSD 2026 Paper Review System
- TSD 2026 CFP (A4 PDF)
- EUSIPCO 2026 Workshop SPQT
- eusipco2026.org submissions
- EMNLP 2026 Call for Main Conference Papers
- ACL ARR CFP
- ACL ARR Authors Guidelines
- TEANGA: Call for papers on corpus linguistics in Irish
- TEANGA Announcements
- BUCC 2026: Building and Using Comparable Corpora
- LaTELL 2026 - Important Dates
- LaTELL-2026.pdf
- VarDial 2026
- VarDial 2026 - Call for Papers
- VarDial 2025 - Program
- NLP4MusA 2026
- LoRes-LM: Second Workshop on Language Models for Low-Resource Languages
- JSALT 2025 Plenary Lectures
- From Seeing it to Experiencing it: Voice Bias - CHI 2026
- From Seeing it to Experiencing it (ACM DL)
- PoliticalNLP 2024
- Speech Synthesis Workshop 2025 (already covered)
- LREC Submission Guidelines
- RESOURCEFUL-2025
Journals
- IEEE Signal Processing Letters
- IEEE Transactions on Audio, Speech and Language Processing
- IEEE Journal of Selected Topics in Signal Processing
- IEEE Open Journal of Signal Processing
- IEEE/ACM TASLP (ACM DL)
- Journal on Audio, Speech, and Music Processing
- EURASIP
- EURASIP Journals
- EURASIP Open Library
- Signal Processing - ScienceDirect
- Speech Communication - ScienceDirect
- Computer Speech & Language - ScienceDirect
- The Journal of the Acoustical Society of America
- JASA Machine Learning in Acoustics collection
- TACL Submission Guidelines
- Transactions of the ACL (TACL)
- The Thai Universal Dependency Treebank (TACL)
- jtei journal Issue 14 (TEI)
- Stanford Agentic Reviewer
- Hear Me Out: Interactive evaluation platform for speech AI
- Hear-Me-Out/src/moshi.py
- International Journal of Natural Language Computing (IJNLC)
Personal Work
KTH
- Speech, Music and Hearing (TMH) - KTH
- tmh/gpu-admin (KTH internal)
- Swedish Research Council: Calls and decisions
- Swedish Research Council: Grant for accessibility to infrastructure
- Swedish Research Council: Research project grant (educational sciences)
- Swedish Research Council: Starting grant (natural and engineering sciences)
- Digital Futures Calls
- Optimizing ASR Models with Semantic Information (KTH DiVA)
- Optimizing ASR Models with Semantic Information (Springer)
- High-accuracy prediction of mental health scores from BERT (KTH DiVA)
- Inter-language Transfer Learning for Visual Speech Recognition (ACL)
- Investigating NMT for Low-Resource: Bavarian (ACL)
Parliamentary & Text Corpora
- CLARIN Resource Families
- ParlaMint: Parliamentary Corpora - CLARIN
- CLARIN:EL - MultiEURLEX
- Sign-Hub WP 2.4 - ORTOLANG
- PFC - Phonologie du Français Contemporain - ORTOLANG
- Spontaneous Dialogues in L1 English - ORTOLANG
- The CLES corpus of spontaneous L2 English - ORTOLANG
- LIDILEM/plspp (GitLab)
- Kategori:Sveriges riksdag (sv Wikipedia)
- jtei/4133 (XML/TEI)
LibriVox & Project Gutenberg
- Wuthering Heights - Project Gutenberg
- Wuthering Heights (dramatic reading) - LibriVox
- Works read by Jason Mills - LibriVox
- Works read by Annie Coleman Rothenberg - LibriVox
- The Brothers Karamazov - Project Gutenberg
- Wigilja Bożego Narodzenia - LibriVox
- Wigilja Bożego Narodzenia - Wikiźródła
- WolneLektury: Dickens Opowieść wigilijna
- Pan Tadeusz - Wikisource
- Konrad Wallenrod (Mickiewicz) - Internet Archive
- Little Women (dramatic reading) - LibriVox
- Little Women - Project Gutenberg
- Robinson Crusoe in Words of One Syllable - Project Gutenberg
- Works by Daniel Defoe - LibriVox
- Recordings of Books on the Ambleside List 2 - Librivox wiki
- Hallowe’en - LibriVox
- The Book of Hallowe’en - Project Gutenberg
- A Visit from St. Nicholas - Wikipedia
- Peter Piper’s Practical Principles - LibriVox
- Peter Piper’s Practical Principles - Project Gutenberg
- ‘Tis Pity She’s a Whore - LibriVox
- If - LibriVox (Kipling)
- Æsop’s Fables - Project Gutenberg
- The Dismissed - LibriVox
- The Maltese Falcon (1930)/Chapter 1 - Wikisource
- Steppenwolf/Preface - Wikisource
- The Good Soldier: Schweik/Book 1/Chapter 1
- Project Gutenberg Open Audiobook Collection
- ntfy notification service
- binwiederhier/ntfy-android
Entertainment & Media
- Panel Show Weekly Schedule - 22 March 2026 : r/panelshow
- Kongen Befaler S12 E10 (with English subs) : r/panelshow
- Bäst i Test (Taskmaster Sweden) S11E03 w/ Eng subs : r/panelshow
- Series 22 - Google Drive
- White Lotus • HBO Max
- The Fish Doorbell
- Midrift - Tell Me Everything (YouTube)
- Midrift - Spotify
- Puscifer: The Story Behind The Coolest Record of 2026 (YouTube)
- Why Is This The Most METAL Scale? (YouTube)
- Uncovering a conspiracy (YouTube)
- SCARY MOVIE 6 Official Trailer (2026) (YouTube)
- Self Defense Expert Answers Self Defense Questions - WIRED (YouTube)
- Sexuality Professor Answers Dating Questions - WIRED (YouTube)
- Why High Masking Autistics Struggle with Transitions (YouTube)
- 5 Must-Know GRINDCORE Riffs (Part 2) (YouTube)
- Clair de lune - Debussy (guitare) (YouTube)
- Tonttu Toljeranti jakso 18 (YouTube)
- Sexuality Professor Answers Dating Questions - LADbible (YouTube)
-
5 Signs You’re A High-Masking Autistic With ADHD
- Difficulty with social interactions
- Intense focus on special interests
- Sensitivity to sensory stimuli
- Adherence to routines
- Executive function challenges
Maths & Foundations
- Essence of linear algebra - YouTube playlist (3Blue1Brown)
- 3Blue1Brown: The Essence of Calculus
- But what is a neural network? - Deep learning chapter 1 (YouTube)
- Transformers, the tech behind LLMs - Deep Learning Chapter 5 (YouTube)
- Let’s build GPT from scratch (Karpathy, YouTube)
- MIT 18.06: The Column Space of A (YouTube)
- Calculus for Mathematicians, Physicists, and Computer Scientists
- Attention Mechanism from Scratch - MachineLearningMastery
- Bahdanau Attention Mechanism - MachineLearningMastery
- The Roadmap of Mathematics for Machine Learning
- eriklindernoren/ML-From-Scratch
- Implementing the Fourier Transform Numerically in Python
- Cover tree - Wikipedia
- Jfeatherstone/CoverTree: Python implementation of cover trees
Misc
- nyrahealth/CrisperWhisper: Verbatim ASR with improved word-level timestamps
- jiaaro/pydub
- hmmlearn GMMHMM API
- anymalign.limsi.fr
- xenon-middleware/xenon/.zenodo.json
- CarlosHolivan/SelfSimilarityMatrices
- UnsupSeg/next_frame_classifier.py
- SegFeat/dicts/phoneme_types.txt
- 2025.fieldmatters-1.1.pdf (ACL)
- Enchrony - Wiley Cognitive Science
- ByT5 documentation (HuggingFace Transformers)
- mlx-whisper · PyPI
- mlm_from_scratch/attn_numpy_demo.py
- Code an LLM from Scratch (freeCodeCamp)
- cisean/barcode_database: Irish grocery barcodes
- tiiuae/Falcon-Perception
- Falcon-Perception/demo/ocr.ipynb
- diva-portal.org/smash/get/diva2:1271586/FULLTEXT01.pdf
- uu.diva-portal.org FULLTEXT01.pdf
- content (dspace.ut.ee)
- dspinellis/speak (already in TTS section)
- SIMH - Wikipedia
- Open SIMH Project
- CDC 1700 - Wikipedia
- Bitsavers: CDC 1700
- Bitsavers: SGI Iris
- open-simh/simtools
- Rhialto/macro11
- SIMH Software Kits
- Preserving code: Zork I, II, III go Open Source
- ZILF
- historicalsource/zork1
- LMAAFY — Let Me Ask AI For You
- Wendy Elvira-García (UB Phonetics Lab)
- Taja Kuzman Pungeršek (Notion)
- webrecorder/pywb: Python Web Archiving Toolkit
- RetroAssembly
- cliopatria.swi-prolog.org
- Gromadzenie POPC PDF (Polish)
- IncusOS: Immutable Linux with ZFS for Containers
- Previously-unknown ancient language discovered in Turkey
- copy/v86
- MiniMax organization card
- hotdog-not-hotdog
- xuanjihe/speech-emotion-recognition
- ace-step/ACE-Step — ACE-Step: A Step Towards Music Generation Foundation Model
- plexusone/omnivoice-core — Voice abstraction layer supporting TTS, STT, and Voice Agents across multiple providers and transport protocols.
There was a 'Moved Permanently' error fetching URL: 'https://x.com/Pseudo_Sid26/status/1992586014460903741'
There was a 'Moved Permanently' error fetching URL: 'https://x.com/TheGlobalMinima/status/1992186189051531567'
There was a 'Moved Permanently' error fetching URL: 'https://x.com/_ARahim_/status/2041979482043744347'
ARahim3/mlx-tune — Fine-tune LLMs on your Mac with Apple Silicon
There was a 'Moved Permanently' error fetching URL: 'https://x.com/GTL094144/status/2041065414659285444'
Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
There was a 'Moved Permanently' error fetching URL: 'https://x.com/ModelScope2022/status/2041343097498374564'
VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning, modelscope, github
- 30 languages: Arabic, Burmese, Chinese, Danish, Dutch, English, Finnish, French, German, Greek, Hebrew, Hindi, Indonesian, Italian, Japanese, Khmer, Korean, Lao, Malay, Norwegian, Polish, Portuguese, Russian, Spanish, Swahili, Swedish, Tagalog, Thai, Turkish, Vietnamese
- Voice design:
wav = model.generate(text="(A young woman, gentle and sweet voice)Hello, welcome to VoxCPM2!", cfg_value=2.0, inference_timesteps=10) - Controllable cloning:
text="(slightly faster, cheerful tone)This is a cloned voice with style control.", - 48kHz output (AudioVAE V2 super-resolution)
- RTF 0.3 (RTX 4090) or ~0.13: Nano-VLLM
- Fine-tuning: “VoxCPM2 supports both full SFT and LoRA fine-tuning with as little as 5–10 minutes of audio”
There was a 'Moved Permanently' error fetching URL: 'https://x.com/awagents/status/2043671651364090125'
There was a 'Moved Permanently' error fetching URL: 'https://x.com/i/trending/2043726207376609641'