Navid MoazzezNavid Moazzez

22 best media & audio GitHub repos to use

The tools behind video, audio and images. Transcription, text to speech, encoding, animation and the libraries every creative pipeline eventually depends on.

Navid Moazzezby Navid Moazzez·Updated Sept 30, 2026·18 min read

Find GitHub repos worth using by topic, language, license and how active they are. Navid's picks come first, and each one opens its own page.

This is a list of the best media & audio GitHub repos.

In fact, it has 22 of them, with Navid's picks first.

So if you want media & audio GitHub repos worth your time, you'll love this list.

The tools behind video, audio and images. Transcription, text to speech, encoding, animation and the libraries every creative pipeline eventually depends on.

Here's what's inside:

Each one comes with what it covers and who it's for.

What are the best media & audio GitHub repos?

Here's the list at a glance.

Owner
OpenAI
Stars
★ 109k
Owner
comfy-org
Stars
★ 134k
Owner
harry0703
Stars
★ 125k
Owner
3b1b
Stars
★ 94k
Owner
FFmpeg
Stars
★ 64k
Owner
calesthio
Stars
★ 61k
Owner
remotion-dev
Stars
★ 60k
Owner
lllyasviel
Stars
★ 53k
Owner
OpenBMB
Stars
★ 38k
Owner
block
Stars
★ 35k
Owner
cjpais
Stars
★ 32k
Owner
anil-matcha
Stars
★ 29k
Owner
SYSTRAN
Stars
★ 25k
Owner
index-tts
Stars
★ 24k
Owner
Wan-Video
Stars
★ 17k
Owner
pipecat-ai
Stars
★ 16k
Owner
duixcom
Stars
★ 16k
Owner
supertone-inc
Stars
★ 14k
Owner
k2-fsa
Stars
★ 14k
Owner
altic-dev
Stars
★ 12k
Owner
GVCLab
Stars
★ 3.8k
Owner
zhouxiaoka
Stars
★ 1.1k

Top 22 media & audio GitHub repos

1. Whisper by OpenAI

Whisper is a general-purpose speech recognition model. It is trained on a large dataset of diverse audio and is also a multitasking model that can perform multilingual speech recognition, speech translation, and language identification.

A Transformer sequence-to-sequence model is trained on various speech processing tasks, including multilingual speech recognition, speech translation, spoken language identification, and voice activity detection. These tasks are jointly represented as a sequence of tokens to be predicted by the decoder, allowing a single model to replace many stages of a traditional speech-processing pipeline. The multitask training format uses a set of special tokens that serve as task specifiers or classification targets.

We used Python 3.9.9 and PyTorch 1.10.1 to train and test our models, but the codebase is expected to be compatible with Python 3.8-3.11 and recent PyTorch versions. The codebase also depends on a few Python packages, most notably OpenAI's tiktoken for their fast tokenizer implementation. You can download and install (or update to) the latest release of Whisper with the following command:

Stars: 109k

Language: Python

License: MIT

Install:

bash brew install ffmpeg 

View Whisper on GitHub · More about Whisper

2. ComfyUI by comfy-org

The most powerful and modular AI engine for content creation.

[![Website][website-shield]][website-url] [![Dynamic JSON Badge][discord-shield]][discord-url] [![Twitter][twitter-shield]][twitter-url] [![Matrix][matrix-shield]][matrix-url]

[![][github-release-shield]][github-release-link] [![][github-release-date-shield]][github-release-link] [![][github-downloads-shield]][github-downloads-link] [![][github-downloads-latest-shield]][github-downloads-link]

Stars: 134k

Language: Python

License: GPL-3.0

Install:

bash pip install -r manager_requirements.txt 

View ComfyUI on GitHub · More about ComfyUI

3. MoneyPrinterTurbo by harry0703

感谢字节火山引擎赞助本项目! 【专属活动优惠】19元Tokens包!享字节自研豆包模型+满血版开源 SOTA模型,覆盖文本、VLM、图像生成,全模态一站配齐:Seed-2.1、Seedream-5.0、GLM-5.2、DeepSeek、Qwen等。不止编程,更能解决 Agent 复杂长程任务 --> 注册即领2500万Tokens,立即前往

感谢 CCSub 赞助本项目!CCSub 是稳定、实惠的 AI API 中转平台,是 Claude Code 官方订阅的超强平替。一个 API Key 即可调用 Claude Opus 4.8、Sonnet 4.6、Haiku 4.5、GPT-5、Gemini 等模型,价格约为官方直连的 1/3,全球直连无需梯子。兼容 Claude Code、Codex、Cursor、Cline、Continue、Windsurf 等所有主流 AI 编程工具。前往 www.ccsub.net 注册即送 $5 体验额度。

感谢 0029.org 云桥 赞助本项目!0029.org 云桥是一个集成了 Claude Code、Codex 以及 Gemini 最新模型的一站式中转平台,为你提供稳定、高效且高性价比的 AI 中转服务。本站提供灵活的包月套餐/按量计费计划,国内直连,无需魔法,极速响应。支持个人和企业接入,价格最低为官方 0.12 折。立即访问。

Stars: 125k

Language: Python

License: MIT

Install:

bash git clone https://github.com/harry0703/MoneyPrinterTurbo.git

View MoneyPrinterTurbo on GitHub · More about MoneyPrinterTurbo

4. Manim by 3b1b

Manim is an engine for precise programmatic animations, designed for creating explanatory math videos.

Note, there are two versions of manim. This repository began as a personal project by the author of 3Blue1Brown for the purpose of animating those videos, with video-specific code available here. In 2020 a group of developers forked it into what is now the community edition, with a goal of being more stable, better tested, quicker to respond to community contributions, and all around friendlier to get started with. See this page for more details.

[!Warning] WARNING: These instructions are for ManimGL only. Trying to use these instructions to install Manim Community/manim or instructions there to install this version will cause problems. You should first decide which version you wish to install, then only follow the instructions for your desired version.

Stars: 94k

Language: Python

License: MIT

Install:

bash pip install manimgl 

View Manim on GitHub · More about Manim

5. FFmpeg by FFmpeg

FFmpeg is a collection of libraries and tools to process multimedia content such as audio, video, subtitles and related metadata. libavcodec provides implementation of a wider range of codecs. libavformat implements streaming protocols, container formats and basic I/O access. libavutil includes hashers, decompressors and miscellaneous utility functions. libavfilter provides means to alter decoded audio and video through a directed graph of connected filters. libavdevice provides an abstraction to access capture and playback devices. libswresample implements audio mixing and resampling routines. libswscale implements color conversion and scaling routines. ffmpeg is a command line toolbox to manipulate, convert and stream multimedia content. ffplay is a minimalistic multimedia player. ffprobe is a simple analysis tool to inspect multimedia content. Additional small tools such as aviocat, ismindex and qt-faststart.

The offline documentation is available in the doc/ directory.

The online documentation is available in the main website and in the wiki.

Stars: 64k

Language: C

License: Other

View FFmpeg on GitHub · More about FFmpeg

6. OpenMontage by calesthio

Bloome lets multiple AI agents (Claude, ChatGPT, DeepSeek, and more) collaborate in one conversation for agentic video pipelines. It has zero setup, runs in the cloud, works on web and mobile, and lets you share a configured agent with your whole team. Try Bloome.

Atlas Cloud is a full-modal AI inference platform that gives developers a single AI API for video generation, image generation, and LLM APIs. Instead of managing multiple vendor integrations, you connect once and get unified access to 300+ curated models across all modalities. Check out Atlas Cloud's new coding plan promotion for more budget-friendly API access.

Turn your AI coding assistant into a full video production studio. Describe what you want in plain language, your agent handles research, scripting, asset generation, editing, and final composition.

Stars: 61k

Language: Python

License: AGPL-3.0

Install:

bash git clone https://github.com/calesthio/OpenMontage.git

View OpenMontage on GitHub · More about OpenMontage

7. Remotion by remotion-dev

Video tools for the agent era. Make videos agentically: Turn your idea into a video using your coding agent. Make videos interactively: Edit and animate using drag and drop. Make videos programmatically: Connect to data, and manage complexity with code.

React Code is the source of truth. Switch your workflow at any point. Design systems: Create a library of animated assets for your organization. Batch rendering: Render millions of videos on your own infrastructure. Applications: Publish a simple tool or a complex video editor.

to get started. Otherwise, read the installation page in the documentation.

Stars: 60k

Language: TypeScript

License: Other

Install:

bash npx create-video@latest 

View Remotion on GitHub · More about Remotion

8. Fooocus by lllyasviel

Fooocus presents a rethinking of image generator designs. The software is offline, open source, and free, while at the same time, similar to many online image generators like Midjourney, the manual tweaking is not needed, and users only need to focus on the prompts and images. Fooocus has also simplified the installation: between pressing "download" and generating the first image, the number of needed mouse clicks is strictly limited to less than 3. Minimal GPU memory requirement is 4GB (Nvidia).

Recently many fake websites exist on Google when you search “fooocus”. Do not trust those, here is the only official source of Fooocus.

The Fooocus project, built entirely on the Stable Diffusion XL architecture, is now in a state of limited long-term support (LTS) with bug fixes only. As the existing functionalities are considered as nearly free of programmartic issues (Thanks to [mashb1t's huge efforts), future updates will focus exclusively on addressing any bugs that may arise.

Stars: 53k

Language: Python

License: GPL-3.0

View Fooocus on GitHub · More about Fooocus

9. VoxCPM by OpenBMB

VoxCPM is a tokenizer-free Text-to-Speech system that directly generates continuous speech representations via an end-to-end diffusion autoregressive architecture, bypassing discrete tokenization to achieve highly natural and expressive synthesis.

VoxCPM2 is the latest major release, a 2B parameter model trained on over 2 million hours of multilingual speech data, now supporting 30 languages, Voice Design, Controllable Voice Cloning, and 48kHz studio-quality audio output. Built on a MiniCPM-4 backbone. 30-Language Multilingual, Input text in any of the 30 supported languages and synthesize directly, no language tag needed Voice Design, Create a brand-new voice from a natural-language description alone (gender, age, tone, emotion, pace …), no reference audio required Controllable Cloning, Clone any voice from a short reference clip, with optional style guidance to steer emotion, pace, and expression while preserving the original timbre Ultimate Cloning, Reproduce every vocal nuance: provide both reference audio and its transcript, and the model continues seamlessly from the reference, faithfully preserving every vocal detail, timbre, rhythm, emotion, and style (same as VoxCPM1.5) 48kHz High-Quality Audio, Accepts 16kHz reference audio and directly outputs 48kHz studio-quality audio via AudioVAE V2's asymmetric encode/decode design, with built-in super-resolution, no external upsampler needed Context-Aware Synthesis, Automatically infers appropriate prosody and expressiveness from text content Real-Time Streaming, RTF as low as ~0.3 on NVIDIA RTX 4090, and ~0.13 accelerated by Nano-vLLM or vLLM-Omni, official vLLM omni-modal serving for VoxCPM2 with PagedAttention and an OpenAI-compatible API Fully Open-Source & Commercial-Ready, Weights and code released under the Apache-2.0 license, free for commercial use

Chinese Dialect: 四川话, 粤语, 吴语, 东北话, 河南话, 陕西话, 山东话, 天津话, 闽南话 [2026.04] We release VoxCPM2, 2B, 30 languages, Voice Design & Controllable Voice Cloning, 48kHz audio output! Weights Docs Playground Technical Report [2025.12] Open-source VoxCPM1.5 weights with SFT & LoRA fine-tuning. ( #1 GitHub Trending) [2025.09] Release VoxCPM Technical Report. [2025.09] Open-source VoxCPM-0.5B weights ( #1 HuggingFace Trending)

Stars: 38k

Language: Python

License: Apache-2.0

Install:

bash pip install voxcpm 

View VoxCPM on GitHub · More about VoxCPM

10. Buzz by block

A workspace where humans and agents build together, on a relay you own.

Vision · Sovereign · Forge · Agents · Architecture · Releasing · Apache 2.0

Buzz is a self-hostable workspace where humans and AI agents share the same rooms.

Stars: 35k

Language: Rust

License: Apache-2.0

Install:

bash git clone https://github.com/block/buzz.git && cd buzz 

View Buzz on GitHub · More about Buzz

11. Handy by cjpais

A free, open source, and extensible speech-to-text application that works completely offline.

Handy is a cross-platform desktop application that provides simple, privacy-focused speech transcription. Press a shortcut, speak, and have your words appear in any text field. This happens on your own computer without sending any information to the cloud.

Handy was created to fill the gap for a truly open source, extensible speech-to-text tool. As stated on handy.computer: Free: Accessibility tooling belongs in everyone's hands, not behind a paywall Open Source: Together we can build further. Extend Handy for yourself and contribute to something bigger Private: Your voice stays on your computer. Get transcriptions without sending audio to the cloud Simple: One tool, one job. Transcribe what you say and put it into a text box

Stars: 32k

Language: Rust

License: MIT

View Handy on GitHub · More about Handy

12. Open Generative Ai by anil-matcha

[](https://muapi.ai?utmsource=github&utmmedium=badge&utmcampaign=open-generative-ai)

The free, open-source alternative to AI Video Platforms. Generate AI images and videos using 400+ state-of-the-art models across 14 studios, no content filters, no closed ecosystem, no subscription fees.

Want to launch this as your own branded AI studio and charge your own customers for it? MuAPI White Label lets you spin up a fully white-labeled version of this app, your logo, your colors, your custom domain, your own pricing, with zero infra to manage. You keep the markup on every generation; MuAPI handles the models, the queue, and the billing plumbing underneath. Your branding, logo, color theme, and a custom domain (e.g. studio.yourbrand.com) Your pricing, set your own credit/subscription prices for end users, keep the margin No infra, no servers, workers, or model hosting to run yourself All studios included, Image, Video, Audio, Lip Sync, Cinema, Workflows, and more, depending on plan

Stars: 29k

Language: JavaScript

License: MIT

Install:

bash npm run electron:build:linux 

View Open Generative Ai on GitHub · More about Open Generative Ai

13. Faster Whisper by SYSTRAN

faster-whisper is a reimplementation of OpenAI's Whisper model using CTranslate2, which is a fast inference engine for Transformer models.

This implementation is up to 4 times faster than openai/whisper for the same accuracy while using less memory. The efficiency can be further improved with 8-bit quantization on both CPU and GPU.

For reference, here's the time and memory usage that are required to transcribe 13 minutes of audio using different implementations: openai/whisper@v20240930 whisper.cpp@v1.7.2 transformers@v4.46.3 faster-whisper@v1.1.0

Stars: 25k

Language: Python

License: MIT

Install:

bash pip install nvidia-cublas-cu12 nvidia-cudnn-cu12==9.* 

View Faster Whisper on GitHub · More about Faster Whisper

14. Index Tts by index-tts

IndexTTS is a zero-shot text-to-speech system that clones a voice from a single reference audio clip. The latest release, IndexTTS-2.5, supports Chinese, English, Japanese, Spanish and Arabic, with fine-grained emotion control, speaking speed control, pronunciation control (Pinyin / CMU phonemes / Japanese Kana), and faster inference than IndexTTS-2.

Example audio files are downloaded on demand from HuggingFace/ModelScope the first time the WebUI starts, so Git LFS is no longer required.

We use uv to manage the project's dependency environment. It is required for a reliable installation:

Stars: 24k

Language: Python

License: Other

Install:

bash git clone https://github.com/index-tts/index-tts.git && cd index-tts 

View Index Tts on GitHub · More about Index Tts

15. Wan 2.1 by Wan-Video

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its performance is even comparable to some closed-source models. Multiple Tasks: Wan2.1 excels in Text-to-Video, Image-to-Video, Video Editing, Text-to-Image, and Video-to-Audio, advancing the field of video generation. Visual Text Generation: Wan2.1 is the first video model capable of generating both Chinese and English text, featuring robust text generation that enhances its practical applications. Powerful Video VAE: Wan-VAE delivers exceptional efficiency and performance, encoding and decoding 1080P videos of any length while preserving temporal information, making it an ideal foundation for video and image generation. May 14, 2025: We introduce Wan2.1 VACE, an all-in-one model for video creation and editing, along with its inference code, weights, and technical report! Apr 17, 2025: We introduce Wan2.1 FLF2V with its inference code and weights! Mar 21, 2025: We are excited to announce the release of the Wan2.1 technical report. We welcome discussions and feedback! Mar 3, 2025: Wan2.1's T2V and I2V have been integrated into Diffusers (T2V I2V). Feel free to give it a try! Feb 27, 2025: Wan2.1 has been integrated into ComfyUI. Enjoy! Feb 25, 2025: We've released the inference code and weights of Wan2.1.

If your work has improved Wan2.1 and you would like more people to see it, please inform us. Helios, a breakthrough video generation model base on Wan2.1 that achieves minute-scale, high-quality video synthesis at 19.5 FPS on a single H100 GPU (about 10 FPS on a single Ascend NPU), without relying on conventional long video anti-drifting strategies or standard video acceleration techniques. Visit their webpage for more details. Video-As-Prompt, the first unified semantic-controlled video generation model based on Wan2.1-14B-I2V with a Mixture-of-Transformers architecture and in-context controls (e.g., concept, style, motion, camera). Refer to the project page for more examples. LightX2V, a lightweight and efficient video generation framework that integrates Wan2.1 and Wan2.2, supports multiple engineering acceleration techniques for fast inference, which can run on RTX 5090 and RTX 4060 (8GB VRAM). DriVerse, an autonomous driving world model based on Wan2.1-14B-I2V, generates future driving videos conditioned on any scene frame and given trajectory. Refer to the project page for more examples. Training-Free-WAN-Editing, built on Wan2.1-T2V-1.3B, allows training-free video editing with image-based training-free methods, such as FlowEdit and FlowAlign. Wan-Move, accepted to NeurIPS 2025, a framework that brings Wan2.1-I2V-14B to SOTA fine-grained, point-level motion control! Refer to their project page for more information. EchoShot, a native multi-shot portrait video generation model based on Wan2.1-T2V-1.3B, allows generation of multiple video clips featuring the same character as well as highly flexible content controllability. Refer to their project page for more information. AniCrafter, a human-centric animation model based on Wan2.1-14B-I2V, controls the Video Diffusion Models with 3DGS Avatars to insert and animate anyone into any scene following given motion sequences. Refer to the project page for more examples. HyperMotion, a human image animation framework based on Wan2.1, addresses the challenge of generating complex human body motions in pose-guided animation. Refer to their website for more examples. MagicTryOn, a video virtual try-on framework built upon Wan2.1-14B-I2V, addresses the limitations of existing models in expressing garment details and maintaining dynamic stability during human motion. Refer to their website for more examples. ATI, built on Wan2.1-I2V-14B, is a trajectory-based motion-control framework that unifies object, local, and camera movements in video generation. Refer to their website for more examples. Phantom has developed a unified video generation framework for single and multi-subject references based on both Wan2.1-T2V-1.3B and Wan2.1-T2V-14B. Please refer to their examples. UniAnimate-DiT, based on Wan2.1-14B-I2V, has trained a Human image animation model and has open-sourced the inference and training code. Feel free to enjoy it! CFG-Zero enhances Wan2.1 (covering both T2V and I2V models) from the perspective of CFG. TeaCache now supports Wan2.1 acceleration, capable of increasing speed by approximately 2x. Feel free to give it a try! DiffSynth-Studio provides more support for Wan2.1, including video-to-video, FP8 quantization, VRAM optimization, LoRA training, and more. Please refer to their examples. Wan2.1 Text-to-Video [x] Multi-GPU Inference code of the 14B and 1.3B models [x] Checkpoints of the 14B and 1.3B models [x] Gradio demo [x] ComfyUI integration [x] Diffusers integration [ ] Diffusers + Multi-GPU Inference Wan2.1 Image-to-Video [x] Multi-GPU Inference code of the 14B model [x] Checkpoints of the 14B model [x] Gradio demo [x] ComfyUI integration [x] Diffusers integration [ ] Diffusers + Multi-GPU Inference Wan2.1 First-Last-Frame-to-Video [x] Multi-GPU Inference code of the 14B model [x] Checkpoints of the 14B model [x] Gradio demo [ ] ComfyUI integration [ ] Diffusers integration [ ] Diffusers + Multi-GPU Inference Wan2.1 VACE [x] Multi-GPU Inference code of the 14B and 1.3B models [x] Checkpoints of the 14B and 1.3B models [x] Gradio demo [x] ComfyUI integration [ ] Diffusers integration [ ] Diffusers + Multi-GPU Inference

Models Download Link Notes, T2V-14B Huggingface ModelScope Supports both 480P and 720P I2V-14B-720P Huggingface ModelScope Supports 720P I2V-14B-480P Huggingface ModelScope Supports 480P T2V-1.3B Huggingface ModelScope Supports 480P FLF2V-14B Huggingface ModelScope Supports 720P VACE-1.3B Huggingface ModelScope Supports 480P VACE-14B Huggingface ModelScope Supports both 480P and 720P

Stars: 17k

Language: Python

License: Apache-2.0

Install:

bash git clone https://github.com/Wan-Video/Wan2.1.git

View Wan 2.1 on GitHub · More about Wan 2.1

16. Pipecat by pipecat-ai

Pipecat is an open-source Python framework for building real-time voice and multimodal conversational agents. Build a single voice agent or a full multi-agent system where specialists hand off, fan out in parallel, and coordinate over a shared bus, locally or distributed across processes and machines. Orchestrate audio and video, AI services, transports, and conversation pipelines effortlessly, so you can focus on what makes your agents unique.

Want to dive right in? Run pipecat init quickstart or follow the quickstart guide. Voice Assistants, natural, streaming conversations with AI Multi-Agent Systems, specialists that hand off, fan out in parallel, or run as sidecars over a shared bus AI Companions, coaches, meeting assistants, characters Multimodal Interfaces, voice, video, images, and more Interactive Storytelling, creative tools with generative media Business Agents, customer intake, support bots, guided flows Complex Dialog Systems, design logic with structured conversations Voice-first: Integrates speech recognition, text-to-speech, and conversation handling Pluggable: Supports many AI services and tools Composable Pipelines: Build complex behavior from modular components Multi-Agent Ready: Each pipeline is an agent. Compose them with handoff, parallel fan-out, sidecar workers, or distributed deployments Real-Time: Ultra-low latency interaction with different transports (e.g. WebSockets or WebRTC)

Building client applications? You can connect to Pipecat from any platform using our official SDKs:

Stars: 16k

Language: Python

License: BSD-2-Clause

Install:

bash git clone https://github.com/pipecat-ai/pipecat.git

View Pipecat on GitHub · More about Pipecat

17. Duix Avatar by duixcom

  1. What's Duix.Avatar 2. Introduction 3. How to Run Locally 4. Open APIs 5. What's New 6. FAQ 7. How to Interact in real time 8. Contact 9. License 10. Acknowledgments 11. Star History

Duix.Avatar is a free and open-source AI avatar project developed by Duix.com.

Seven years ago, a group of young pioneers chose an unconventional technical path, developing a method to train digital human models using real-person video data. Unlike traditional costly 3D digital human approaches, we leveraged AI-generated technology to create ultra-realistic digital humans, slashing production costs from hundreds of thousands of dollars to just $1,000. This innovation has empowered over 10,000 enterprises and generated over 500,000 personalized avatars for professionals across fields, educators, content creators, legal experts, medical practitioners, and entrepreneurs, dramatically enhancing their video production efficiency. However, our vision extends beyond commercial applications. We believe this transformative technology should be accessible to everyone. To democratize digital human creation, we've open-sourced our cloning technology and video production framework. Our commitment remains: breaking down technological barriers to make cutting-edge tools available to all. Now, anyone with a computer can freely craft their own AI Avatar and produce videos at zero cost, this is the essence of Duix.Avatar.

Stars: 16k

Language: C

License: Other

Install:

bash docker-compose -f docker-compose-linux.yml up -d 

View Duix Avatar on GitHub · More about Duix Avatar

18. Supertonic by supertone-inc

This repository will be archived, and there will be no further development or official support for the open-source Supertonic models.

Voice Builder will no longer be accessible after August 31, 2026.

For details about the service changes, timeline, and information for existing Voice Builder users, please see the official announcement.

Stars: 14k

Language: Swift

License: MIT

Install:

bash pip install supertonic 

View Supertonic on GitHub · More about Supertonic

19. OmniVoice by k2-fsa

OmniVoice is a state-of-the-art massively multilingual zero-shot text-to-speech (TTS) model supporting over 600 languages. Built on a novel diffusion language model-style architecture, it generates high-quality speech with superior inference speed, supporting voice cloning and voice design.

We recommend using a fresh virtual environment (e.g., conda, venv, etc.) to avoid conflicts.

Intel Arc GPUs (Alchemist and Battlemage architectures) are supported via PyTorch's XPU backend.

Stars: 14k

Language: Python

License: Apache-2.0

Install:

bash pip install torch==2.8.0 torchaudio==2.8.0 

View OmniVoice on GitHub · More about OmniVoice

20. FluidVoice by altic-dev

Open source voice-to-text dictation app for macOS with on-device AI enhancement.

[!NOTE] FluidVoice is on macOS today. iOS and Windows are on the way, join the waitlist to get notified when we launch: altic.dev/fluid/waitlist

[!IMPORTANT] This project is free and open source under GPLv3. If FluidVoice is useful to you, please star the repository, it helps visibility and keeps development going.

Stars: 12k

Language: Swift

License: GPL-3.0

Install:

bash brew install --cask fluidvoice 

View FluidVoice on GitHub · More about FluidVoice

21. PersonaLive by GVCLab

1 University of Macau    2 Dzine.ai    3 GVC Lab, Great Bay University

   [ ] If you find PersonaLive useful or interesting, please give us a Star! Your support drives us to keep improving. [ ] Fix bugs (If you encounter any issues, please feel free to open an issue or contact me! ) [x] [2026.05.15] Release training code. [x] [2026.02.21] PersonaLive is accepted by CVPR2026. [x] [2025.12.29] Enhance WebUI (Support reference image replacement). [x] [2025.12.22] Supported streaming strategy in offline inference to generate long videos on 12GB VRAM! [x] [2025.12.17] ComfyUI-PersonaLive is now supported! (Thanks to @okdalto) [x] [2025.12.15] Release paper! [x] [2025.12.12] Release inference code, config, and pretrained weights! [x] This project is released for academic research only. [x] Users must not use this repository to generate harmful, defamatory, or illegal content. [x] The authors bear no responsibility for any misuse or legal consequences arising from the use of this tool. [x] By using this code, you agree that you are solely responsible for any content generated.

We present PersonaLive, a real-time and streamable diffusion framework capable of generating infinite-length portrait animations.

Stars: 3.8k

Language: Python

License: Apache-2.0

Install:

bash git clone https://github.com/GVCLab/PersonaLive

View PersonaLive on GitHub · More about PersonaLive

22. AutoClip by zhouxiaoka

An intelligent video clipping and collection recommendation system based on AI, supporting automatic Bilibili video download, subtitle extraction, intelligent slicing, and collection generation. Features Quick Start Project Structure Configuration User Guide Development Guide FAQ Changelog License Contributing Contact Intelligent Video Clipping: AI-powered video content analysis for high-quality automatic clipping Bilibili Video Download: Support for automatic Bilibili video download and subtitle extraction Smart Collection Recommendations: AI automatically analyzes slice content and recommends related collections Manual Collection Editing: Support drag-and-drop sorting, adding/removing slices One-Click Package Download: Support one-click package download for all slices and collections Modern Web Interface: React + TypeScript + Ant Design Real-time Processing Status: Real-time display of processing progress and logs Python 3.8+ Node.js 16+ DashScope API Key or SiliconFlow API Key (for AI analysis) Docker 20.10+ Docker Compose 2.0+ DashScope API Key or SiliconFlow API Key (for AI analysis)

One-click deployment, no complex environment setup required!

Configure your API keys in data/settings.json. You can choose between DashScope and SiliconFlow APIs: DashScope: Visit Alibaba Cloud Console AI Services Tongyi Qianwen API Key Management SiliconFlow: Visit SiliconCloud Login API Keys Create New API Key

Stars: 1.1k

Language: Python

License: MIT

Install:

bash pip install -r requirements.txt 

View AutoClip on GitHub · More about AutoClip

Final thoughts on media & audio GitHub repos

The best media & audio repo is the one you come back to. Pick 2 or 3 from this list, give them a month, and keep the ones you look forward to.

Navid Moazzez

AI business strategist & AI OS builder

Navid Moazzez helps creators and founders master AI and build their own AI Operating System (AI OS) to automate their business and life.

Navid.me is reader-supported. When you buy through links on this site, I may earn an affiliate commission. Learn more.

More GitHub repos

More free tools

Related MCP servers & CLIs

Free AI newsletter

The most actionable AI newsletter for founders

Every week, get proven AI strategies, curated tools, and step-by-step systems to grow your audience, create better content, and build a profitable creator business.

No fluff, no filler, no BS. Just five minutes each week that might level up your online business and life.

P.S. Sign up now to get free access to my ultimate AI tools guide for creators.

Loved by 10,000+ readers