Claude sorted tabs, 22/03/2026
Misc. interesting things.
- Parliamentary Corpora
- Grapheme-to-Phoneme (G2P)
- Northern Sami
- Multilingual / Low-Resource ASR
- CTC Decoding and Forced Alignment
- Disentangled Speech Representations / Voice Conversion
- Distributed Training
- LibriVox and English Accent Data
- OCR / HTR Tools
- Query-by-Example Spoken Term Detection (QbE-STD)
- Speech Embeddings and MSEB
- Speech Codecs and Audio Tokenization
- WFST for ASR
- Universal Dependencies and Spoken Language Parsing
- RDF / Linked Data
- Swedish Radio Minority Language Programming
- Web Archiving
- Irish Language / Irish ASR
- Hungarian Language Learning
- Sentence alignment tools
- Anki tools
Parliamentary Corpora
ParlaMint / ParlaCLARIN
- ParlaMint: Comparable and Interoperable Parliamentary Corpora (CLARIN)
- ParlaMint I and II project information
- Multilingual comparable corpora of parliamentary debates ParlaMint 5.0 (CLARIN.SI)
- clarin-eric/ParlaMint GitHub
- Parla-CLARIN TEI Schema for Corpora of Parliamentary Proceedings
- clarin-eric/parla-clarin: Schema for modelling parliamentary debates
- ParlaCLARIN IV Workshop on Creating, Analysing, and Increasing Accessibility of Parliamentary Corpora
Swerik / Riksdagen
- swerik-project repositories
- The Swedish Parliament Corpus 1867 – 2022 (LREC 2024 paper)
- swerik-project/riksdagen-records
- swerik-project/riksdagen-motions
- swerik-project/riksdagen-persons: Metadata on politicians
- swerik-project/pyparlaclarin: Python package for Parla-CLARIN XML
- swerik-project/riksdagen-volumeG-alto: Alto files for OCRd volume G
- corpus-walkthrough.ipynb (Colab)
-
[Parliamentary documents, laws and legislative history Sveriges riksdag](https://www.riksdagen.se/en/contact-and-visit/the-riksdag-library/parliamentary-documents-laws-and-legislative-history/)
Other corpora / projects
- ParlaSpeech - Parliamentary Speech Corpus
- How to Use the ParlaSpeech Corpus Through a Concordancer
- The siParl corpus of Slovene parliamentary proceedings
- KBLab/rixvox-v2 (Hugging Face)
- Hansard - UK Parliament
- ajdapretnar/AI-perspectives: AI in British and Slovenian parliament
Related papers
- Encoding Interruptions in Parliamentary Data: From Applause to Interjections and Laughter
- Qualitative Comparison of Native and Machine-Translated Parliamentary Debates
-
[Voices of the Parliament Modern Languages Open](https://modernlanguagesopen.org/articles/10.3828/mlo.v0i0.295)
CLARIN resource families
- Parliamentary Corpora
- Legal Corpora
- Sign Language Resources
- Licenses and CLARIN Categories
- CLARIN Licensing Framework
- JRC EU DGT Translation Memory Parsebank DGT-UD 1.0
Grapheme-to-Phoneme (G2P)
Tools / models
- spring-media/DeepPhonemizer: Grapheme to phoneme with deep learning
- DeepPhonemizer training notebook (Colab)
- lingjzhu/CharsiuG2P: Multilingual G2P in 100 languages
- charsiu/g2p_multilingual_byT5_small_100
- CharsiuG2P fine-tuning low-resource notebook
- fdemelo/g2p-mbyt5-12l-ipa-childes-espeak
- rhasspy/wiktionary2dict: Extract IPA pronunciations from Wiktionary XML dump
- open-dict-data/ipa-dict: Monolingual wordlists with IPA pronunciation
- CUNY-CL/wikipron (phoneme data)
Papers
- ByT5 model for massively multilingual G2P (2204.03067)
- PolyIPA: Multilingual Phoneme-to-Grapheme Conversion (2412.09102)
- SoundChoice: G2P Models with Semantic Disambiguation (Interspeech 2022)
- Transformer Based G2P Conversion (Interspeech 2019)
- T5G2P: Using Text-to-Text Transfer Transformer for G2P (Interspeech 2021)
- G2P using LSTM recurrent neural networks (IEEE ICASSP 2015)
- Massively Multilingual Neural G2P Conversion (ACL 2017)
- Joint-sequence models for G2P conversion (ScienceDirect)
- Jointly learning to align and convert graphemes to phonemes (ASRU 2017)
- Italianising English words with G2P in TTS
- Stress Assignment in Letter to Sound Rules
- Phoneme-to-Grapheme Conversion Based Large-Scale Pre-Training for ASR
- Knowledge of language origin improves pronunciation of proper names
- Improving Proper Name Recognition by Adding Learned Pronunciation Variants (LREC 2010)
- Galescu 2002 (ICSLP) - joint grapheme-phoneme alignment
Pronunciation lexicons (Scandinavian)
- NB Pronunciation (Norsk ordbank) - Språkbanken
- NB Uttale Work on Dialects (PDF)
- NST uttaleleksikon for dansk - Språkbanken
- Lingit uttaleleksikon for nynorsk - Språkbanken
- Norsk ordbank - bokmål 2005 - Språkbanken
- Norsk ordbank - nynorsk 2012 - Språkbanken
Pronunciation dictionaries / data
- flexthink/librig2p-nostress-space (Hugging Face)
- australian-lexicon: Australian IPA phonemes
- GlobalPhone Language Models
- CMU Sphinx Acoustic and Language Models (SourceForge)
Northern Sami
Corpora
- Giellagas corpus (Kielipankki)
- Giellagas license page
- Kielipankki access
- LIA sápmi - LIA corpus of Sami dialects (Språkbanken)
- Northern Sami YleAreena Subtitle Corpus (Zenodo)
- Northern Saami interactive text corpus (Giellatekno)
- COMEDI editor (Clarino)
- UD_North_Sami-Giella
ASR models
- GetmanY1/wav2vec2-large-sami-22k-finetuned
- Fine-Tune W2V2-Bert for low-resource ASR
- facebook/w2v-bert-2.0
- SeamlessM4T-v2 docs
- facebook/seamless-m4t-v2-large
- divvun-tts/multi-sami HF Space
- Getman 2024 Interspeech paper
- anyspeech/zipa-large-crctc-ns-800k
Phonetics / phonology
- Guide to North Sámi Pronunciation (Oahpa Muinna)
- Forvo: Northern Sami pronunciation dictionary
- Forvo: čáhci pronunciation
- Wiktionnaire: Annexe:Prononciation/same du Nord
- Wayback: Odden Saami phonology PDF
- Sami Grammar - Vocabulary (archived)
- The acoustic manifestation of consonant gradation in Northern Sami (KTH)
- An analysis of North Saami gradation
- Dialectal variation of duration patterns in Finnmark North Sámi (KTH)
- Initially and Finally Stressed Vowel Sequences of Guovdageaidnu Dialect (KTH)
- Helsinki paper (content link)
Learning resources / language documentation
- DigiSami project
- DigiSami IT paper
- Yle: Say it in Saami soundboard
- Learn North-Sámi part 1 (YouTube)
- About the Sámi languages (YouTube)
- Finnish-Samish YouTube channel
- Northern sami - 10 common words (YouTube)
- Numbers in Northern Sami 1-100 (YouTube)
- From Language to Language From Mind to Nation 9 (ovttas.no)
- About the Sámi Parliament (Sametinget)
- Radio and Television Archive - Kavi (Finland)
Multilingual / Low-Resource ASR
Omnilingual ASR (Facebook Research)
- facebookresearch/omnilingual-asr: Open-Source Multilingual Speech Recognition for 1600+ Languages
- facebook/omniASR-LLM-7B (Hugging Face)
- omnilingual-asr data preparation README
- omnilingual-asr per-language results (7B LLM ASR)
OWSM (Open Whisper-style Speech Models)
Audio LLM / speech-language model frameworks
- X-LANCE/SLAM-LLM: Speech, Language, Audio, Music with LLM
- SLAM-LLM speech-to-speech README
- FunAudioLLM/CosyVoice: Multilingual large voice generation model
- modelscope/FunASR
- Continuous Audio Language Models (2509.06926)
- Bridging the Modality Gap: Softly Discretizing Audio for LLM-based ASR (2506.05706)
- Retrieval Augmented Generation based context discovery for ASR (2509.19567)
- NLE: Non-autoregressive LLM-based ASR by Transcript Editing
- Bridging the gap: Speech-LLM and end-to-end for multilingual conversational ASR
W2V2 / multilingual fine-tuning
- Fine-Tune W2V2-Bert for low-resource ASR (HF blog)
- facebook/w2v-bert-2.0 (Hugging Face)
- facebook/seamless-m4t-v2-large (Hugging Face)
- GetmanY1/wav2vec2-large-sami-22k-finetuned (Hugging Face)
- Getman 2024 Interspeech paper
- KBLab/kb-whisper-large (Hugging Face)
SIGUL: Low-resource and endangered language NLP
- SIGUL 2024 Proceedings (LREC-COLING)
- Assessing Pre-Built Speaker Recognition Models for Endangered Language Data
- Tandem Long-Short Duration-based Modeling for ASR
- SIGUL 2026 Joint Workshop CfP
- LaTeLL 2026: Language Technologies for Low-resource Languages
CTC Decoding and Forced Alignment
CTC decoding
- CTC Networks and Language Models: Prefix Beam Search Explained (Medium)
- kensho-technologies/pyctcdecode: Fast and lightweight CTC beam search decoder
- torchaudio ctc_decoder documentation
- torchaudio CTCDecoder documentation
- ASR Inference with CTC Decoder tutorial (torchaudio)
- ASR inference with CUDA CTC decoder tutorial
- SpeechBrain CTC prefix beam search docs
- FlexCTC: GPU-powered CTC Beam Decoding With Advanced Contextual Abilities
Forced alignment tools
- lumaku/ctc-segmentation: Segment audio and obtain utterance alignments
- m-bain/whisperX alignment module
- tabahi/bournemouth-forced-aligner
- BFA: Real-time Multilingual Text-to-speech Forced Alignment (2509.23147)
- tabahi/contexless-phonemes-CUPE
- CUPE: Contextless Universal Phoneme Encoder (2508.15316)
- mlx-audio Qwen3 forced aligner
- mlx-community/Qwen3-ForcedAligner-0.6B-4bit (Hugging Face)
- bertsky/nmalign: Forced alignment of string lists by fuzzy string matching
- emilyahn/align_cs
Phoneme recognition / IPA models
- kgnlp/allophant: Multilingual phoneme recognizer with zero-shot generalization
- lingjzhu/zipa: IPA-based ASR
- anyspeech/zipa-large-crctc-ns-800k (Hugging Face)
- PRiSM: Benchmarking Phone Realization in Speech Models (2601.14046)
- lingjzhu/clap-ipa
- anyspeech/ipapack_plus_meta (Hugging Face)
- allophoible
- xinjli/ucla-phonetic-corpus
- VinAIResearch/XPhoneBERT (Interspeech 2023)
Segmentation / other
- felixkreuk/UnsupSeg: Unsupervised segmentation
- Dysfluent WFST: Zero-Shot Speech Dysfluency Transcription (2505.16351)
Disentangled Speech Representations / Voice Conversion
Key models / implementations
- Berkeley-Speech-Group/sparc: Speech Articulatory Coding
- SPARC demo notebook (Colab)
- cheoljun95/Speech-Articulatory-Coding (Hugging Face model)
- Berkeley-Speech-Group/sylber: Syllabic Embedding Representation of Speech
- light1726/SpeechTripleNet: End-to-End Disentangled Speech (content, timbre, prosody)
- SpeechTripleNet demo page
- auspicious3000/SpeechSplit: Unsupervised Speech Decomposition via Triple Information Bottleneck
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers (ICML 2022)
- YoungSeng/SRD-VC
- synspeech.github.io: Learning Disentangled Speech Representations
Papers
- Learning Disentangled Speech Representations (2311.03389)
- Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval (2401.08096)
- VQ-CL: Learning Disentangled Speech with Contrastive Learning and Vector Quantization (ICASSP 2023)
- SpeechTripleNet: End-to-End Disentangled Speech Representation Learning (ACM MM 2023)
- Unsupervised speech decomposition via triple information bottleneck (ICML 2020)
- Adversarially Learning Disentangled Speech Representations for Robust Multi-Factor Voice Conversion (Interspeech 2021)
- Disentangled Speech Representation for Cross-Lingual Voice Conversion Using β-VAE (IEEE 2023)
- An Overview of Voice Conversion and its Challenges (2008.03648)
- ContentVec: An Improved Self-Supervised Speech Representation by Disentangling Speakers (2204.09224)
- MeanVoiceFlow: One-step Nonparallel Voice Conversion with Mean Flows (2602.18104)
- Speech to Speech Synthesis for Voice Impersonation (2602.16721)
- Representation Learning with Contrastive Predictive Coding (1807.03748)
- beta-VAE: Learning Basic Visual Concepts with a Constrained Variational Framework (OpenReview)
Pitch / prosody tools
- maxrmorrison/torchcrepe: PyTorch CREPE pitch tracker
- interactiveaudiolab/penn: Pitch Estimating Neural Networks
- maxrmorrison/torbi: Viterbi decoding in PyTorch
- marl/crepe
Datasets
- ylacombe/expresso (Hugging Face)
- EXPRESSO dataset demo page
- EXPRESSO paper (Interspeech 2023, 2308.05725)
Distributed Training
Ray
- ray-project/ray: AI compute engine and distributed runtime
- Ray Train: Scalable Model Training
- Ray Train Overview
- Get Started with Distributed Training using PyTorch Lightning
- Get Started with Distributed Training using Hugging Face Transformers
- Loading Data in Ray
Submitit (Slurm)
Lhotse
LibriVox and English Accent Data
LibriVox wiki / data collection
- LibriVox wiki main page
- Accents Table - LibriVox wiki
- Recordings of Books on the Ambleside List
- Recordings of Books on the Ambleside List 2
- Other British Readers on LibriVox (RuthieG’s CataBlog)
Multi-accent speech datasets
- Open-source Multi-speaker Corpora of English Accents in the British Isles (LREC 2020)
- The Edinburgh International Accents of English Corpus
- The Voice Bank Corpus: Large Regional Accent Speech Database (IEEE 2013)
- CSTR VCTK Corpus: English Multi-speaker Corpus
- KBLab/rixvox-v2 (Swedish) (Hugging Face)
- Google Nigerian English (MFA models)
- kth-tmh/google-britain-ireland (Hugging Face)
- LibriTTS British Accents (GitHub)
- Speech Accent Archive (Weinberger & Kunath)
- cvaccents notebook (Kathy Reid)
- Effect AI Scripted Speech 1.0 - English (Mozilla Data Collective)
- Common Voice Spontaneous Speech 3.0 - English
- Common Voice Spontaneous Speech 3.0 - Irish
- IDEA: Accents and Dialects of Europe
- Accents of English (Cambridge book)
Edinburgh / CSTR datasets
- Centre for Speech Technology Research (CSTR) data
- NST (Natural Speech Technology) data (Edinburgh)
- Listening test materials for modern speech synthesis evaluation
OCR / HTR Tools
Document processing platforms
- Arkindex: a document processing platform (Teklia)
- Arkindex GitLab
- Automatic Text Recognition / PyLaia (Teklia GitLab)
- Explicit language modeling with n-grams (PyLaia docs)
- SCRIBE project
Models
- ibm-granite/granite-docling-258M (Hugging Face)
- granite-docling-258M demo Space
- HTRflow: A tool for HTR and OCR (HF blog)
- Supercharge your OCR Pipelines with Open Models (HF blog)
- Grounded Fine-tuning notebook (smol-vision)
- deepseek-ai/DeepSeek-OCR: Contexts Optical Compression
Query-by-Example Spoken Term Detection (QbE-STD)
Tools / implementations
- idiap/CNN_QbE_STD: CNN based Query by Example Spoken Term Detection
- anupsingh15/BEST-STD2.0: Token subword mapping
- Yushi-Hu/Query-by-Example
- Yushi-Hu/Acoustic-Span-Embeddings
- rainmaker29/SpokenWord2Vec
Datasets / benchmarks
- QUESST 2014 Multilingual Database for QbE Keyword Spotting (BUT Speech@FIT)
- MediaEval 2014 Workshop Proceedings (CEUR-WS Vol-1263)
- MediaEval 2013 SWS - Spoken Web Search
- MediaEval 2013 Workshop Proceedings (CEUR-WS Vol-1043)
- MediaEval 2012 Workshop Proceedings (CEUR-WS Vol-927)
Papers
- H-QuEST: Accelerating QbE STD with Hierarchical Indexing (2506.16751)
- CNN Based Query by Example Spoken Term Detection (Interspeech 2018)
- Attention-Based Audio Embeddings for Query-by-Example (2210.08624)
- Query-by-Example Spoken Term Detection using Attentive Pooling Networks (IEEE 2020)
- Query-by-example spoken term detection using bottleneck features and HMM (IEEE 2016)
- Spoken Word2Vec: Learning Skipgram Embeddings from Speech (2311.09319)
- Vocabulary independent spoken term detection (ACM SIGIR 2007)
- A lattice-based approach to QbE spoken document retrieval (ACM SIGIR 2008)
- Unsupervised Pattern Discovery in Speech (IEEE 2008)
- A Nonparametric Bayesian Approach to Acoustic Model Discovery (ACL 2012)
- Phonetic-and-semantic embedding of spoken words (Interspeech 2024)
- QbE-STD using bottleneck features (2410.04091)
- An introduction to voice search (IEEE 2008)
Related: Speech-to-Retrieval
- Speech-to-Retrieval (S2R): A new approach to voice search (Google Research)
- google/svq dataset (Hugging Face)
Speech Embeddings and MSEB
Massive Sound Embedding Benchmark (MSEB)
- MSEB project page (Google Research)
- MSEB paper (2602.07143)
- google-research/mseb GitHub
- MSEB encoder module
- MSEB segmentation encoder
- MSEB whisper encoder
- MSEB raw encoder
- From Waveforms to Wisdom: The New Benchmark for Auditory Intelligence (Google blog)
- NeurIPS 2025 MSEB poster
Speech-to-Retrieval (S2R)
Google speech embedding (older)
- google-research/speech_embedding
- speech_embedding record_train.ipynb
- speech_embedding speech_commands.ipynb
CLAP
Speech Codecs and Audio Tokenization
Models / implementations
- ZhangXInFD/SpeechTokenizer: Unified Speech Tokenizer for Speech LMs (ICLR 2024)
- OpenMOSS-Team/SpeechTokenizer (Hugging Face)
- 0nutation/USLM: Unified Speech Language Model
- 0nutation/SLMTokBench: SpeechTokenizer benchmark
- novateur/WavTokenizer-large-speech-75token (Hugging Face)
- novateur/WavTokenizer-medium-music-audio-75token (Hugging Face)
- neuphonic/neucodec: 50hz, 0.8kbps, 24kHz audio codec
- neuphonic/neucodec (Hugging Face)
- neuphonic/emilia-yodas-english-neucodec dataset (Hugging Face)
- modelscope/FunCodec
- yangdongchao/AcademiCodec
- ktvoice/Codec (Hugging Face)
- speechbrain/hifigan-wavlm-k1000-LibriTTS (Hugging Face)
SODA (Scaling Open Discrete Audio)
Papers / surveys
- Discrete Audio Tokens: More Than a Survey! (2506.10274)
- SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models (OpenReview ICLR 2024)
- DeCodec: Rethinking Audio Codecs as Universal Disentangled Representation Learners (2509.09201)
- Scaling Open Discrete Audio Foundation Models with Interleaved Tokens (2602.16687)
WFST for ASR
System combination / multi-stream decoding
- A WFST Framework for Single-Pass Multi-Stream Decoding (Interspeech 2016)
- Combination of multiple aligned recognition outputs using WFST and LSTM (IEEE 2015)
- Language Model Combination and Adaptation Using WFSTs (NASA NTRS)
- WFST Enabled Solutions to ASR Problems: Beyond HMM Decoding (Google Research)
Dysfluency / other applications
- Dysfluent WFST: A Framework for Zero-Shot Speech Dysfluency Transcription and Detection (2505.16351)
- DysfluentWFST IPA2CMU config
- transducersaurus/regex2wfst.py
Universal Dependencies and Spoken Language Parsing
UD tools and annotation
- Universal Dependencies main site
- UD tools page
- DepEdit - tool for manipulating dependency trees
- nats/TrUDucer (GitLab)
- furkan7258/boat: Boğaziçi University Annotation Tool
- LR-POR/cl-conllu: Tool for working with CoNLL-U files in CL
Spoken / non-standard UD
- ufcompling/spoken_parsing (MaChAmp predict)
- stavros-bompolas/ngud-transformer: NGUD transformer
- xiulinyang/UD-CHILDES
- UniversalDependencies/UD_English-CHILDES
- UD_Yiddish-YiTB
- Swedish UD
- UniversalDependencies/UD_North_Sami-Giella
NLP tools
- stanfordnlp/stanza: Retrain UD models
- nlp-uoregon/trankit: Light-Weight Transformer-based Multilingual NLP
- TurkuNLP/Turku-paraphrase-corpus (Swedish)
- Surface-Syntactic UD data
- heinz-jeffrey/bufia: Formal language learning from positive examples
Workshops
- UDW26 - Universal Dependencies Workshop
- 2025.udw-1.9.pdf (ACL Anthology)
- 2025.udw-1.5.pdf (ACL Anthology)
- 2025.udw-1.6.pdf (ACL Anthology)
- The status of function words in UD (Glossa)
WFST for morphology / phonology
- Learning Cross-Dialectal Morphophonology with Syllable Structure Constraints (VarDial 2025)
- transducersaurus/regex2wfst.py
RDF / Linked Data
Standards and specs
Tools / frameworks
- sparql-generate/sparql-generate: SPARQL-Generate over Apache Jena
- ClioPatria: SWI-Prolog Semantic Web Server
- kba/jsonld-rapper: Create RDF from JSON-LD with rapper
- figshare/cc0rdfhosting: CC0 RDF Hosting
- RDF workflow for figshare (proof of concept, figshare)
WordNet / linguistic linked data
- jrvosse/wordnet-3.0-rdf: Princeton WordNet 3.0 as linked open data
- jmccrae/wordnet-angular: Princeton WordNet Interface
- VU WordNet dataset (datahub.io)
Abgeleitete Textformate / derived text formats
-
[Abgeleitete Textformate: Text und Data Mining mit urheberrechtlich geschützten Textbeständen ZfdG](https://zfdg.de/2020_006)
Swedish Radio Minority Language Programming
Sveriges Radio channels
- Sameradion
- Sameradiopodden (all episodes)
- Tablå Sameradion
- Meänraatio
- Radio Romano
- Terni Generatcia (romani och svenska)
- Mitt jiddischorakel
- Jiddisch far alle
- News in other languages
- All podcasts/programs
SVT (Swedish Television)
-
[Jiddisch/Yiddish SVT Play](https://www.svtplay.se/kategori/jiddisch-yiddish)
Sveriges Radio Open API
- This is Swedish Radio’s open API
- RSS and Podcasting info
- API channels list
- RSS feed: Sameradion program 2327
- RSS feed: pod 4901
- RSS feed: program 2054
- Now playing API (channel 213)
Archive
Web Archiving
Tools / frameworks
- Webrecorder: Web Archiving for All
- webrecorder/pywb: Core Python Web Archiving Toolkit
- pywb documentation
- pywb recording mode configuration
- pywb Memento API
- ReplayWeb.page (Webrecorder)
Warchaeology (Norwegian National Library)
Formats
Irish Language / Irish ASR
ASR systems and papers
- Automatic Speech Recognition for Irish: the ABAIR-ÉIST System (CLTW 2022)
- Fotheidil: an Automatic Transcription System for the Irish Language (CLTW 2025)
- Cross-dialect lexicon optimisation for an endangered language ASR system: the case of Irish
- Automatic Speech Recognition for Irish: testing lexicons and language models (IEEE 2022)
Celtic Language Technology Workshop (CLTW)
- Proceedings of the 4th Celtic Language Technology Workshop (LREC 2022)
- Second Celtic Language Technology Workshop (LATTICE)
Irish language NLP
- Kevin Scannell: Tiomsú Corpais don Taighde Foclóireachta (CFG2020)
- Kevin Scannell: Neural Models for Predicting Celtic Mutations
- Kevin Scannell: Claonadh Inscne, Líonraí Néaracha, agus an Ghaeilge
-
[Irish corpus from the web Sketch Engine](https://www.sketchengine.eu/gatenten-irish-corpus/) - Innovative Irish Language and AI projects at DCU
- Call for papers: TEANGA special edition on Irish-language corpus linguistics
Irish language text resources
- Leigh Leat
- Stór Scéalta - Leigh Leat
- ClubLeabhar.com: na leabhair
- Sources for Connemara Irish - Gaeilge Chonamara
- Bailiúchán Béaloidis Árann (Aran Folklore Collection)
Irish phonology
- A contribution to the phonology of Desi-Irish (Wikisource)
- UCLA Phonetics: Gaelic, Irish
- UCLA Phonetics: Word List for Gaelic, Irish (Galway dialect)
Hungarian Language Learning
Grammar references
- HungarianReference.com - Verb conjugation overview
- The may/permission suffix: -hat/het
- The conditional mood: would/should
- Expressing need with kell + personal infinitives
- Future tense and using two verbs
- Repeated actions with -gat/get
- Splitting coverbs from verb root
- Syntax, word order and forming sentences
- The word ‘hogy’ and its uses
- Hungarian Dative case: -nak/-nek
- HungarianReference.com links/dictionary/resources
- Hungarian verbs - Wikipedia
- Magyar helyesírás – Wikipédia
- Hungarian Grammar/Vocabulary - Wikibooks
Courses / learning resources
- The MagyarOK textbook series
- Easy Hungarian - Bits of Hungarian
- Best Resources for Learning Hungarian (Catch Budapest)
- Hungarian resources - Lindie Botes
- r/hungarian wiki: Learn Hungarian Online
- Fluent Forever base vocabulary list
- How to say “IN” in Hungarian: -ban, -ben, -n, -on (YouTube)
- Bogyó és Babóca - Hungarian children’s show (YouTube)
- Zsenileszek channel (YouTube)
- Harisnyás Pippi (Pippi Longstocking in Hungarian, Dailymotion)
Gutenberg texts (Hungarian)
- Books: Language: Hungarian - Project Gutenberg
- Világ ura (Master of the World) - Jules Verne
- Kárpáthy Zoltán - Mór Jókai
- Utazás a Holdba (From the Earth to the Moon) - Jules Verne
- Figurák - Géza Gárdonyi
- Az emberiség képviselői (Representative Men) - Emerson
- Népvilág - Mór Jókai
- Légy jó mindhalálig - Zsigmond Móricz
- Virradóra - Mór Jókai
- Huckleberry Finn kalandjai - Mark Twain
- Életbölcseség (Aphorisms on the Wisdom of Life) - Schopenhauer
- Im-ígyen szóla Zarathustra (Thus Spoke Zarathustra) - Nietzsche
- Grimm Fairy Tales (Hungarian)
- Books by Jókai, Mór - Project Gutenberg
Sentence alignment tools
- hunalign – sentence aligner (Hungarian NLP)
- loomchild/maligna: Bilingual sentence aligner
- pombredanne/anymalign: Multilingual aligner for SMT