Transformers over molecular strings: SMILES and SELFIES encoders, de novo generators, text-to-molecule translation, and reaction prediction.
will PRO
wrice
AI & ML interests
Interested in the applications of generative models for Speech Synthesis and NLP.
Recent Activity
liked a model 6 days ago
DBD-research-group/AudioProtoPNet-5-BirdSet-XCL updated a collection 6 days ago
SMILES & SELFIES Language Models updated a collection 6 days ago
SMILES & SELFIES Language ModelsOrganizations
Full-Duplex Speech Models
Conversational models that listen and speak simultaneously, including voice control, interactivity tuning, and multimodal streaming.
Automatic Speech Recognition
ASR models for multilingual transcription and streaming speech recognition, with notes on use cases and licenses.
-
Qwen/Qwen3-ASR-1.7B
Automatic Speech Recognition • 2B • Updated • 1.09M • • 1.16k -
nvidia/parakeet-tdt-0.6b-v3
Automatic Speech Recognition • 0.6B • Updated • 613k • • 1.18k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.26M • • 3.43k -
kyutai/stt-1b-en_fr
Automatic Speech Recognition • 1.0B • Updated • 137
Biomedical Text Embedding
Super Resolution
Image & video super-resolution / upscaling models: GAN, transformer, diffusion, and latent upscalers.
Text to Speech
TTS models spanning lightweight synthesis, voice cloning, multilingual speech, and streaming generation.
Voice Activity Detection
Speech/non-speech segmentation models: ASR end-pointing, diarization pre-processing, streaming turn detection. Server-side and on-device/browser.
SMILES & SELFIES Language Models
Transformers over molecular strings: SMILES and SELFIES encoders, de novo generators, text-to-molecule translation, and reaction prediction.
Super Resolution
Image & video super-resolution / upscaling models: GAN, transformer, diffusion, and latent upscalers.
Full-Duplex Speech Models
Conversational models that listen and speak simultaneously, including voice control, interactivity tuning, and multimodal streaming.
Text to Speech
TTS models spanning lightweight synthesis, voice cloning, multilingual speech, and streaming generation.
Automatic Speech Recognition
ASR models for multilingual transcription and streaming speech recognition, with notes on use cases and licenses.
-
Qwen/Qwen3-ASR-1.7B
Automatic Speech Recognition • 2B • Updated • 1.09M • • 1.16k -
nvidia/parakeet-tdt-0.6b-v3
Automatic Speech Recognition • 0.6B • Updated • 613k • • 1.18k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.26M • • 3.43k -
kyutai/stt-1b-en_fr
Automatic Speech Recognition • 1.0B • Updated • 137
Voice Activity Detection
Speech/non-speech segmentation models: ASR end-pointing, diarization pre-processing, streaming turn detection. Server-side and on-device/browser.
Biomedical Text Embedding