Interesting links, 04/01/2024
Misc. interesting things.
Dao-AILab/flash-attention — Fast and memory-efficient exact attention
facebookincubator/velox — A C++ vectorized database acceleration library aimed to optimizing query engines and data processing systems.
How 🤗 Accelerate runs very large models thanks to PyTorch
karkirowle/relative_phoneme_analysis — Repository for phoneme analysis on word-level Kaldi/ESPNet ASR transcripts
prajdabre/yanmtt — Yet Another Neural Machine Translation Toolkit
google-research-datasets/TextNormalizationCoveringGrammars — Covering grammars for English and Russian text normalization
WavJourney: Compositional Audio Creation with Large Language Models
Minimum Word Error Rate Training with Language Model Fusion for End-to-End Speech Recognition
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/din0s_/status/1742235150530851120'
VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
thuhcsi/VAENAR-TTS — The official implementation of VAENAR-TTS, a VAE based non-autoregressive TTS model.
Automatic Generation of Subtitles for Videos of the Government of La Rioja
The Properly Illustrated Transformer
Efficient Sequence Transduction by Jointly Predicting Tokens and Durations
Instant3D: Instant Text-to-3D Generation
LRM: Large Reconstruction Model for Single Image to 3D
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/begusgasper/status/1655981693516517378'
Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks
‘Dair’ (Live @ Urban Assault 2018)
lingjzhu/CharsiuG2P — Multilingual G2P in 100 languages
kNN-CTC: Enhancing ASR via Retrieval of CTC Pseudo Labels, “code”
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/lateinteraction/status/1737578879454425202'
Speculative Decoding for 2x Faster Whisper Inference
SHI-Labs/VCoder — VCoder: Versatile Vision Encoders for Multimodal Large Language Models, arXiv 2023
ConvNets Match Vision Transformers at Scale
SD-HuBERT: Self-Distillation Induces Syllabic Organization in HuBERT
An Introduction to Transformers
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/MattNiessner/status/1724795600456310844'
Improving Large-scale Deep Biasing with Phoneme Features and Text-only Data in Streaming Transducer
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/reach_vb/status/1726699176698732929'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/reach_vb/status/1727065880918409674'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/sophiamyang/status/1733505991600148892'
wellecks/ntptutorial — Tutorial on neural theorem proving
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/rasbt/status/1734234160154185730'
THE LITTLE BOOK OF DEEP LEARNING
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/_philschmid/status/1734992933764411788'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/jxmnop/status/1734961947227897983'
open-mmlab/Amphion — Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
SpeechAct: Towards Generating Whole-body Motion from Speech
Fine-tuning Whisper for Dutch Language: The Crucial Role of Size
Introduction to Speech Processing
OML-Team/open-metric-learning — Library for metric learning pipelines and models.
haotian-liu/LLaVA — [NeurIPS’23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
Simplifying Transformer Blocks
Advanced RAG Techniques: an Illustrated Overview
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/_akhaliq/status/1742757369895960950'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/DuaneJRich/status/1742777245821989224'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/reach_vb/status/1742261240141918684'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/bclavie/status/1742950315278672040'
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/lateinteraction/status/1736804963760976092'
Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models
LLM Augmented LLMs: Expanding Capabilities through Composition
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/jerryjliu0/status/1743077679258320925'
Mathematical Introduction to Deep Learning: Methods, Implementations, and Theory
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/_philschmid/status/1742888388401811795'
This AI Paper from Meta Introduces Hyper-VolTran: A Novel Neural Network for Transformative 3D Reconstruction and Rendering, paper
Phi-2: The surprising power of small language models
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
MotionScript: Natural Language Descriptions for Expressive 3D Human Motions
pjyazdian/Gesture2Vec — This is an official PyTorch implementation of “Gesture2Vec: Clustering Gestures using Representation Learning Methods for Co-speech Gesture Generation” (IROS 2022).
neuromorphs/NIR — Neuromorphic Intermediate Representation reference implementation
PEFT for Speech: Unveiling Optimal Placement, Merging Strategies, and Ensemble Techniques
What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANs
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/lateinteraction/status/1743009556521975893'
100 tiny changes to transform your life: from the one-minute rule to pyjama yoga
The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank Reduction
Decoder-only Architecture for Speech Recognition with CTC Prompts and Text Data Augmentation
LiteLlama: Reduced-Scale Llama — We present an open-source reproduction of Meta AI’s LLaMa 2. However, with significantly reduced model sizes, LiteLlama-460M-1T has 460M parameters trained with 1T tokens.
Token 1.3: What is Retrieval-Augmented Generation (RAG)?
VikParuchuri/surya — Accurate line-level text detection and recognition (OCR) in any language
gchrupala/neurospoken — Neural models of spoken language - LOT Winter school 2024
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/ChristophMolnar/status/1745731602682982675'
My AI Timelines Have Sped Up (Again)
Mixtral 8x7B is currently the best open-source LLM, surpassing GPT-3.5
Foundations of Vector Retrieval
GARField: Group Anything with Radiance Fields
AlphaGeometry: An Olympiad-level AI system for geometry, code
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/francoisfleuret/status/1748011011590799462'
SceneVerse: Scaling 3D Vision-Language Learning for Grounded Scene Understanding