EvolvingLMMs-Lab/NEO — Native Vision-Language Models

Deta Surf

Audio-to-Image Bird Species Retrieval without Audio-Image Pairs via Text Distillation

SAVGBench: Benchmarking Spatially Aligned Audio-Video Generation

pronounce.voanews.com - Afghanistan

Paza Bench

Top 5 Embedding Models for Your RAG Pipeline

Custom Kernels for All from Codex and Claude

tenstorrent/tt-metal — TT-NN operator library, and TT-Metalium low level kernel programming model.

Baidu just dropped an open-source multimodal AI that it claims beats GPT-5 and Gemini: ERNIE-4.5-VL-28B-A3B-Thinking

5 open-source remote desktop tools prove that nobody should use TeamViewer anymore

Human brain cells on a chip learned to play Doom in a week

WAXAL - A Large-Scale Multilingual African Language Speech Corpus, dataset

rustfs/rustfs — RustFS is an open-source, S3-compatible high-performance object storage system supporting migration and coexistence with other S3-compatible platforms such as MinIO and Ceph.

A neural network for modeling human concept formation, understanding and communication

40,000-year-old signs show humans were recording information long before writing

TW-Sound580K: A Regional Audio-Text Dataset with Verification-Guided Curation for Localized Audio-Language Modeling

Focus Then Listen: An Empirical Study of Plug-and-Play Audio Enhancer for Noise-Robust Large Audio Language Models

Multi-Loss Learning for Speech Emotion Recognition with Energy-Adaptive Mixup and Frame-Level Attention

davanstrien/ocr-bench-britannica-results-qwen35-viewer

Discrete Audio Tokens More Than a Survey!, taxonomy

Sleepwalking/SHIRO — Phoneme-to-speech alignment toolkit based on liblrhsmm

danielcopper/wezterm-session-manager — Lua script enhancement for WezTerm that provides functionality to save, load, and restore terminal sessions

MFA - Phone groups

Speech-Omni-Lite: Portable Speech Interfaces for Vision-Language Models

SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass

MV-RAG: Retrieval Augmented Multiview Diffusion

CaviraOSS/OpenMemory — Local persistent memory store for LLM applications including claude desktop, github copilot, codex, antigravity, etc.

DeepSeek OCR

dreamtheater123/Awesome-SpeechLM-Survey

WiT: Waypoint Diffusion Transformers via Trajectory Conflict Navigation

ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model

CohereLabs

DashengTokenizer: One layer is enough for unified audio understanding and generation, models

Voxtral TTS Demo

microsoft/harrier-oss-v1-27b

T5Gemma-TTS Technical Report

RushOnline/midi2hydrogen

AffectSpeech: A Large-Scale Emotional Speech Dataset with Fine-Grained Textual Descriptions for Speech Emotion Captioning and Synthesis

powertab/powertabeditor

What is the Heilmeier Catechism?

The Snack Sound Toolkit

A Speech Recognition Extension to Snack

chdh/klatt-syn

40,000-year-old signs show humans were recording information long before writing

Java to Kotlin Conversion Comes to Visual Studio Code

zpforlove/AG-REPA — Official code for AG-REPA: Causal Layer Selection for Representation Alignment in Audio Flow Matching (ICML 2026). Unified TTS+TTA Flow Matching with BiT-C / LASP / FoG-A diagnostics.

AustinZhang/AG-REPA — AG-REPA model.

Principles and Practice of Deep Representation Learning

Qwen3.5: Towards Native Multimodal Agents

Zaneham/Booth — Open-source CUDA, Triton and HIP compiler targeting multiple GPU and CPU architectures.

QwenLM/Qwen3-TTS

SSVD-O: Parameter-Efficient Fine-Tuning with Structured SVD for Speech Recognition

Exploring Fine-Tuning of Large Audio Language Models for Spoken Language Understanding under Limited Speech Data

Inverse-Hessian Regularization for Continual Learning in ASR

Libation — A free, open-source application for downloading and managing your Audible audiobooks

Mbucari/AAXClean — Decrypt Audible aax and aaxc files.

Qwen/Qwen3-ASR-1.7B

LESS: Large Language Model Enhanced Semi-Supervised Learning for Speech Foundational Models Using in-the-wild Data

Self-Supervised Speech Models Encode Phonetic Context via Position-dependent Orthogonal Subspaces

[b]=[d]-[t]+[p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic, arXiv, code

RadEar: A Self-Supervised RF Backscatter System for Voice Eavesdropping and Separation

rasbt/llm-architecture-gallery, site

ASK: Adaptive Self-improving Knowledge Framework for Audio Text Retrieval

Claude Code’s creator keeps sharing tips, and they all made my experience better

NikolaiKyhne/RWSAMamba-UNet

Implementing the Fourier Transform Numerically in Python: A Step-by-Step Guide

patrick-kidger/equinox — Elegant easy-to-use neural networks + scientific computing in JAX. https://docs.kidger.site/equinox/

We got Claude to teach open models how to write CUDA kernels!

Voxtral Realtime

Blaizzy/mlx-audio

ML Intern Takes Our Post-Training Internship Test

huggingface/ml-intern — an open-source ML engineer that reads papers, trains models, and ships ML models

TIGER-AI-Lab/OpenResearcher — A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis

OpenMOSS-Team/MOSS-Audio-8B-Thinking

Origin Stories Kneecap

No as a Service

huggingface/nfsserve

systemd/casync — Content Addressable Data Synchronizer

xolox/dedupfs — A Python FUSE file system that features transparent deduplication and compression which make it ideal for archiving backups.

containers/fuse-overlayfs

tree-sitter/tree-sitter

CohereLabs/cohere-transcribe-03-2026

QwenLM/FlashQLA — high-performance linear attention kernel library built on TileLang

tile-ai/tilelang — Domain-specific language designed to streamline the development of high-performance GPU/CPU/Accelerators kernels

French Archaeologist Says He Cracked a Mysterious 4,000-Year-Old Bronze Age Script From Ancient Iran

Representing Biomedical Literature as a Filesystem through Agent-Native Indexing

vLLM: Using Docker

UAF: A Unified Audio Front-end LLM for Full-Duplex Speech Interaction

45 years later, earliest DOS source code transcribed from a stack of old printouts found in a garage — code was open-sourced to mark 86-DOS 1.00’s anniversary

86-DOS_1.00

Original Apollo 11 code open-sourced by NASA — original Command Module and Lunar Module code repos are now public domain resources

chrislgarry/Apollo-11

‘Clean-room reimplementation’ of DR-DOS hits early beta, modernizing the operating system 38 years after its debut — runs Doom, Warcraft, SimCity, and other period-appropriate titles

meichthys/foss_photo_libraries

fathah/hermes-desktop — Desktop Companion for Hermes Agent

nousresearch/hermes-agent

X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning

k2-fsa/OmniVoice — High-Quality Voice Cloning TTS for 600+ Languages

OmniVoice space

Ancient Mesopotamian cuneiform texts reach new audiences through major digital archive

MekongPhon: A Large-Scale Parallel IPA Corpus for Lao and Khmer

A Comprehensive Full-Form Lexicon for Arabic NLP and Speech Technology

Saudi ASWAT: A Large-Scale Corpus of Spontaneous Saudi Arabic Speech

Probing Discrete Speech Tokens of Spoken Language Models

An Enhanced Pipeline for the Manzini-Savoia Dialect Corpus

Evaluating Phonetically Weighted and Unweighted Distance Measures in Dialectometry

Chunkwise Aligners for Streaming Speech Recognition

physics-intern: an autonomous agentic framework for physics research

nagamuslim/novnc-audio-plugin