Interesting links, 02/08/2026
Misc. interesting things.
EvolvingLMMs-Lab/NEO — Native Vision-Language Models
Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation
SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation
pronounce.voanews.com - Afghanistan
Top 5 Embedding Models for Your RAG Pipeline
Custom Kernels for All from Codex and Claude
tenstorrent/tt-metal — TT-NN operator library, and TT-Metalium low level kernel programming model.
Baidu just dropped an open-source multimodal AI that it claims beats GPT-5 and Gemini: ERNIE-4.5-VL-28B-A3B-Thinking
5 open-source remote desktop tools prove that nobody should use TeamViewer anymore
- RustDesk (affero)
- remotely
- MeshCentral
- Apache Guacamole
- DWService
Human brain cells on a chip learned to play Doom in a week
WAXAL - A Large-Scale Multilingual African Language Speech Corpus, dataset
rustfs/rustfs — RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.
A neural network for modeling human concept formation, understanding and communication
40,000-year-old signs show humans were recording information long before writing
davanstrien/ocr-bench-britannica-results-qwen35-viewer
Discrete Audio Tokens More Than a Survey!, taxonomy
Sleepwalking/SHIRO — Phoneme-to-speech alignment toolkit based on liblrhsmm
danielcopper/wezterm-session-manager — Lua script enhancement for WezTerm that provides functionality to save, load, and restore terminal sessions
Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models
SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass
MV-RAG: Retrieval Augmented Multiview Diffusion
CaviraOSS/OpenMemory — Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.
dreamtheater123/Awesome-SpeechLM-Survey
WiT: Waypoint Diffusion Transformers via Trajectory Conflict Navigation
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
DashengTokenizer: One layer is enough for unified audio understanding and generation, models
What is the Heilmeier Catechism?
There was a 'Moved Permanently' error fetching URL: 'https://x.com/hooeem/status/2030720614752039185'
A Speech Recognition Extension to Snack
40,000-year-old signs show humans were recording information long before writing
Java to Kotlin Conversion Comes to Visual Studio Code
zpforlove/AG-REPA — Official code for AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching (ICML 2026). Unified TTS+TTA Flow Matching with BiT-C / LASP / FoG-A diagnostics.
AustinZhang/AG-REPA — AG-REPA model.
Principles and Practice of Deep Representation Learning
Qwen3.5: Towards Native Multimodal Agents
Zaneham/Booth — Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.
SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition
Inverse-Hessian Regularization for Continual Learning in ASR
Libation — A free, open-source application for downloading and managing your Audible audiobooks
Mbucari/AAXClean — Decrypt Audible aax and aaxc files.
Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces
[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic, arXiv, code
RadEar: A Self-Supervised RF Backscatter System for Voice Eavesdropping and Separation
rasbt/llm-architecture-gallery, site
ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval
Claude Code’s creator keeps sharing tips, and they all made my experience better
Implementing the Fourier Transform Numerically in Python: A Step-by-Step Guide
patrick-kidger/equinox — Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/
We got Claude to teach open models how to write CUDA kernels!
ML Intern Takes Our Post-Training Internship Test
huggingface/ml-intern — an open-source ML engineer that reads papers, trains models, and ships ML models
TIGER-AI-Lab/OpenResearcher — A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis
OpenMOSS-Team/MOSS-Audio-8B-Thinking
systemd/casync — Content Addressable Data Synchronizer
xolox/dedupfs — A Python FUSE file system that features transparent deduplication and compression which make it ideal for archiving backups.
CohereLabs/cohere-transcribe-03-2026
QwenLM/FlashQLA — high-performance linear attention kernel library built on TileLang
tile-ai/tilelang — Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels
French Archaeologist Says He Cracked a Mysterious 4,000-Year-Old Bronze Age Script From Ancient Iran
Representing Biomedical Literature as a Filesystem through Agent-Native Indexing
UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction
meichthys/foss_photo_libraries
fathah/hermes-desktop — Desktop Companion for Hermes Agent
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning
k2-fsa/OmniVoice — High-Quality Voice Cloning TTS for 600+ Languages
Ancient Mesopotamian cuneiform texts reach new audiences through major digital archive
MekongPhon: A Large-Scale Parallel IPA Corpus for Lao and Khmer
A Comprehensive Full-Form Lexicon for Arabic NLP and Speech Technology
Saudi ASWAT: A Large-Scale Corpus of Spontaneous Saudi Arabic Speech
Probing Discrete Speech Tokens of Spoken Language Models
An Enhanced Pipeline for the Manzini-Savoia Dialect Corpus
Evaluating Phonetically Weighted and Unweighted Distance Measures in Dialectometry
Chunkwise Aligners for Streaming Speech Recognition
physics-intern: an autonomous agentic framework for physics research