Interesting links, 04/04/2023
Misc. interesting things.
Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation
@misc{bengio2021flow,
title={Flow Network based Generative Models for Non-Iterative Diverse Candidate Generation},
author={Emmanuel Bengio and Moksh Jain and Maksym Korablyov and Doina Precup and Yoshua Bengio},
year={2021},
eprint={2106.04399},
archivePrefix={arXiv},
primaryClass={cs.LG}
}
Make-A-Video: Text-to-Video Generation without Text-Video Data
@misc{singer2022makeavideo,
doi = {10.48550/ARXIV.2209.14792},
url = {https://arxiv.org/abs/2209.14792},
author = {Singer, Uriel and Polyak, Adam and Hayes, Thomas and Yin, Xi and An, Jie and Zhang, Songyang and Hu, Qiyuan and Yang, Harry and Ashual, Oron and Gafni, Oran and Parikh, Devi and Gupta, Sonal and Taigman, Yaniv},
title = {Make-A-Video: Text-to-Video Generation without Text-Video Data},
year = {2022},
}
A real-time filled pause detection system for spontaneous speech recognition
@inproceedings{goto99_eurospeech,
author={Masataka Goto and Katunobu Itou and Satoru Hayamizu},
title={A real-time filled pause detection system for spontaneous speech recognition},
year=1999,
booktitle={Proc. 6th European Conference on Speech Communication and Technology (Eurospeech 1999)},
pages={227--230},
doi={10.21437/Eurospeech.1999-60}
}
TransFusion: Transcribing Speech with Multinomial Diffusion, code – not open source.
@misc{baas2022transfusion,
title={TransFusion: Transcribing Speech with Multinomial Diffusion},
author={Matthew Baas and Kevin Eloff and Herman Kamper},
year={2022},
eprint={2210.07677},
archivePrefix={arXiv},
primaryClass={eess.AS}
}
The Norwegian Parliamentary Speech Corpus
@inproceedings{solberg-ortiz-2022-norwegian,
title = "The {N}orwegian Parliamentary Speech Corpus",
author = "Solberg, Per Erik and
Ortiz, Pablo",
booktitle = "Proceedings of the Thirteenth Language Resources and Evaluation Conference",
month = jun,
year = "2022",
address = "Marseille, France",
publisher = "European Language Resources Association",
pages = "1003--1008",
}
There is more to Hungarian than goulash! Grammar Course for Beginners
sp-nitech/diffsptk — A differential version of SPTK
soobinseo/Transformer-TTS — A Pytorch Implementation of “Neural Speech Synthesis with Transformer Network”
jxzhanggg/nonparaSeq2seqVC_code — Implementation code of non-parallel sequence-to-sequence VC
Many-to-English: Data v2 and Statistics
There was a 'Moved Permanently' error fetching URL: 'https://twitter.com/0xDesigner/status/1642554817590566915'
neonbjb/tortoise-tts — A multi-voice TTS system trained with an emphasis on quality
urschrei/pyzotero — Pyzotero: a Python client for the Zotero API
NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
@misc{tan2022naturalspeech,
title={NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality},
author={Xu Tan and Jiawei Chen and Haohe Liu and Jian Cong and Chen Zhang and Yanqing Liu and Xi Wang and Yichong Leng and Yuanhao Yi and Lei He and Frank Soong and Tao Qin and Sheng Zhao and Tie-Yan Liu},
year={2022},
eprint={2205.04421},
archivePrefix={arXiv},
primaryClass={eess.AS}
}
Self-Instruct: Aligning Language Model with Self Generated Instructions, code
@misc{wang2022selfinstruct,
title={Self-Instruct: Aligning Language Model with Self Generated Instructions},
author={Yizhong Wang and Yeganeh Kordi and Swaroop Mishra and Alisa Liu and Noah A. Smith and Daniel Khashabi and Hannaneh Hajishirzi},
year={2022},
eprint={2212.10560},
archivePrefix={arXiv},
primaryClass={cs.CL}
}