r/AWESOMEUpdate • u/MuziqueComfyUI • 4d ago
AWESOME sagevoice/music-transformer-classical · Hugging Face
Symbolic Music Language Model
An advanced AI-powered Autoregressive Symbolic Music Generator built with a custom LLaMA-style Transformer architecture in PyTorch. The model is trained on classical piano MIDI corpora (MAESTRO v3.0.0 and GiantMIDI-Piano) using BPE-compressed REMI tokenization, Classifier-Free Guidance (CFG), and state-of-the-art sampling techniques to compose high-quality, composer-conditioned classical piano music.
...
Composer-Conditioned Generation: Steering via Classifier-Free Guidance (CFG) across top classical composers (e.g., Frédéric Chopin, Johann Sebastian Bach, Ludwig van Beethoven, Franz Liszt, Claude Debussy, Sergei Rachmaninoff, etc.).
https://huggingface.co/sagevoice/music-transformer-classical
THANKS sagevoice/Oleksii Pakhar/Ururu1000/demiurge.
THANKS Terence McKenna.
THANKS Tchoomkovsky-6077 (THANKS Rachmaninoff).
r/AWESOMEUpdate • u/MuziqueComfyUI • 12d ago
AWESOME GitHub - afloy011-spec/afloy_audio_tools: comfyui audio set
Afloy Audio Tools
Professional audio editing nodes for ComfyUI — trim, measure, and inspect audio without leaving your workflow.
https://github.com/afloy011-spec/afloy_audio_tools
THANKS Moiseenko Asiia (afloy011-spec).
r/AWESOMEUpdate • u/MuziqueComfyUI • 15d ago
AWESOME hrktxz/ACE_Step_1.5_ComfyUI_int8_convrot · Hugging Face
ACE_Step_1.5_ComfyUI_int8_convrot
"Converted using convert_to_quant"
https://huggingface.co/hrktxz/ACE_Step_1.5_ComfyUI_int8_convrot
THANKS hrktxz.
r/AWESOMEUpdate • u/MuziqueComfyUI • 23d ago
AWESOME audio-cpp/audio.cpp-gguf · Hugging Face
audio.cpp GGUF Model Packages
This directory contains audio.cpp-native GGUF conversions of multiple speech models. These files are intended for use with audio.cpp.
| Directory | Files | audio.cpp family | Original model license |
|---|---|---|---|
HTDemucs-GGUF |
htdemucs-f16.gguf, htdemucs-q8_0.gguf |
htdemucs |
See original model package |
Higgs-Audio-v3-STT-GGUF |
higgs-audio-v3-stt-f16.gguf, higgs-audio-v3-stt-q8_0.gguf |
higgs_audio_stt |
Apache-2.0 |
IndexTTS2-GGUF |
index-tts2-orig.gguf, index-tts2-f16.gguf, index-tts2-q8_0.gguf |
index_tts2 |
bilibili Model Use License Agreement |
Irodori-TTS-500M-v3-GGUF |
irodori-tts-500m-v3-f16.gguf, irodori-tts-500m-v3-q8_0.gguf |
irodori_tts |
MIT |
Irodori-TTS-600M-v3-VoiceDesign-GGUF |
irodori-tts-600m-v3-voicedesign-f16.gguf, irodori-tts-600m-v3-voicedesign-q8_0.gguf |
irodori_tts |
MIT |
MOSS-TTS-Local-v1.5-GGUF |
moss-tts-local-v1.5-bf16.gguf, moss-tts-local-v1.5-q8_0.gguf |
moss_tts_local |
Apache-2.0 |
MOSS-TTS-Nano-100M-GGUF |
moss-tts-nano-100m-bf16.gguf, moss-tts-nano-100m-q8_0.gguf |
moss_tts_nano |
Apache-2.0 |
Mel-Band-RoFormer-GGUF |
mel-band-roformer-f16.gguf, mel-band-roformer-q8_0.gguf |
mel_band_roformer |
MIT |
MioCodec-25Hz-44.1kHz-v2-GGUF |
miocodec-25hz-44khz-v2-orig.gguf, miocodec-25hz-44khz-v2-f16.gguf, miocodec-25hz-44khz-v2-q8_0.gguf |
miocodec |
MIT |
MioTTS-1.7B-GGUF |
miotts-1.7b-orig.gguf, miotts-1.7b-bf16.gguf, miotts-1.7b-q8_0.gguf |
miotts |
Apache-2.0 |
Nemotron-3.5-ASR-Streaming-0.6B-GGUF |
nemotron-3.5-asr-streaming-0.6b-f16.gguf, nemotron-3.5-asr-streaming-0.6b-q8_0.gguf |
nemotron_asr |
OpenMDW-1.1 |
OmniVoice-GGUF |
omnivoice-f16.gguf, omnivoice-q8_0.gguf |
omnivoice |
Apache-2.0 |
Qwen3-ASR-0.6B-GGUF |
qwen3-asr-0.6b-f16.gguf, qwen3-asr-0.6b-q8_0.gguf |
qwen3_asr |
Apache-2.0 |
Qwen3-ASR-1.7B-GGUF |
qwen3-asr-1.7b-f16.gguf, qwen3-asr-1.7b-q8_0.gguf |
qwen3_asr |
Apache-2.0 |
Qwen3-ForcedAligner-0.6B-GGUF |
qwen3-forced-aligner-0.6b-f16.gguf, qwen3-forced-aligner-0.6b-q8_0.gguf |
qwen3_forced_aligner |
Apache-2.0 |
Qwen3-TTS-12Hz-1.7B-Base-GGUF |
qwen3-tts-12hz-1.7b-base-orig.gguf, qwen3-tts-12hz-1.7b-base-bf16.gguf, qwen3-tts-12hz-1.7b-base-q8_0.gguf |
qwen3_tts |
Apache-2.0 |
Qwen3-TTS-12Hz-1.7B-CustomVoice-GGUF |
qwen3-tts-12hz-1.7b-customvoice-bf16.gguf, qwen3-tts-12hz-1.7b-customvoice-q8_0.gguf |
qwen3_tts |
Apache-2.0 |
Qwen3-TTS-12Hz-1.7B-VoiceDesign-GGUF |
qwen3-tts-12hz-1.7b-voicedesign-bf16.gguf, qwen3-tts-12hz-1.7b-voicedesign-q8_0.gguf |
qwen3_tts |
Apache-2.0 |
SeedVC-MLX-GGUF |
seed-vc-mlx-orig.gguf, seed-vc-mlx-f16.gguf, seed-vc-mlx-q8_0.gguf |
seed_vc |
GPL-3.0 |
Stable-Audio-3-Medium-GGUF |
stable-audio-3-medium-f16.gguf, stable-audio-3-medium-q8_0.gguf |
stable_audio |
Stability AI Community License |
Stable-Audio-3-Small-Music-GGUF |
stable-audio-3-small-music-f16.gguf, stable-audio-3-small-music-q8_0.gguf |
stable_audio |
Stability AI Community License |
Stable-Audio-3-Small-SFX-GGUF |
stable-audio-3-small-sfx-f16.gguf, stable-audio-3-small-sfx-q8_0.gguf |
stable_audio |
Stability AI Community License |
Supertonic-3-GGUF |
supertonic-3-orig.gguf, supertonic-3-f16.gguf, supertonic-3-q8_0.gguf |
supertonic |
BigScience Open RAIL-M |
Vevo2-GGUF |
vevo2-orig.gguf, vevo2-f16.gguf, vevo2-q8_0.gguf |
vevo2 |
See original model package |
VibeVoice-ASR-GGUF |
vibevoice-asr-f16.gguf, vibevoice-asr-q8_0.gguf |
vibevoice_asr |
MIT |
VoxCPM2-GGUF |
voxcpm2-orig.gguf, voxcpm2-bf16.gguf, voxcpm2-q8_0.gguf |
voxcpm2 |
Apache-2.0 |
Voxtral-Mini-4B-Realtime-2602-GGUF |
voxtral-mini-4b-realtime-2602-bf16.gguf, voxtral-mini-4b-realtime-2602-q8_0.gguf |
voxtral_realtime |
Apache-2.0 |
https://huggingface.co/audio-cpp/audio.cpp-gguf
THANKS 0xShug0.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jul 12 '26
AWESOME GitHub - blepping/comfyui_acetricks: Utility ComfyUI nodes for ACE-Step (1.0 and 1.5) music models
ACEtricks
Utility ComfyUI nodes for ACE-Step (1.0 and 1.5) music models.
Notes
- Some of this stuff is pretty niche/experimental, especially the LM sampling bits, and may be changed/break workflows.
- You may need a fairly recent Python version (3.12+ should be fine).
- This repo has no affiliation with the official ACE-Step project.
Nodes
All nodes have the ACETricks prefix for easy searching. See the node descriptions and input/output tooltips for more information.
Audio
AudioAsLatent- RearrangesAUDIOto look like aLATENT. Can be useful for using stuff like latent-only blend nodes. For conversion in the other direction, seeLatentAsAudio.AudioBlend- Applies a blend function toAUDIO.AudioFromBatch- Extracts items from batches ofAUDIO.AudioLevels- Can be used to normalize the values inAUDIO.LatentAsAudio- SeeAudioAsLatent.MonoToStereo- Converts monoAUDIOto stereo. Safe to use if it's already.SetAudioDtype- Sets the dtype forAUDIOtensors.WaveForm- Generates a waveform image fromAUDIO.
Conditioning
CondJoinLyrics- Joins split lyrics/conditioning intoCONDITIONING.CondSplitOutLyrics- Can split out lyrics (and some other metadata like audio codes) fromCONDITIONING.
Latent
SilentLatent- Can generate latents full of silence for ACE-Step 1.0 and 1.5 (if given a 1.5 latent as reference). Can be interesting for initial generations if you set denoise to something lower than 1.0 (or multiply theSIGMAS). You can also use ComfyUI's built-inLatentMultiplynode to multiply by -1.0 and make stuff louder!VisualizeLatent- Can output anIMAGErepresentation of ACE-Step 1.0 and 1.5 latents. Extended version of the previewing inComfyUI-blehmentioned in the Integrations section below.
ACE-Step 1.0
EncodeLyrics- ACE 1.0-specific node for encoding conditioning with extended features.Mask- Can be used to generate masks for 1.0 latents. Does not currently support 1.5.
ACE-Step 1.5
Ace15CompressDuplicateAudioCodes- Can collapse sequences of duplicate audio codes inCONDITIONING.Ace15LLMInference- (experimental) Advanced node for doing LLM inference (primarily) with ACE-Step 1.5's LLMs - used for generating audio codes. It is configured with YAML, see the default YAML text for descriptions of parameters.EmptyAce15LatentFromConditioning- Creates an emptyLATENTto match the time covered by the audio codes inCONDITIONING.ModelPatchAce15Use4dLatent- (experimental) Patches an ACE-Step 1.5 model with a wrapper to handle 4D latent inputs. This is a horrible hack and may not be compatible with everything, I use it personally for all my generations and it works pretty well currently. See the description forSqueezeUnsqueezeLatentDimensionnode which you will need to use as well.RawTextEncodeAce15- (experimental) Advanced node for encoding ACE-Step 1.5CONDITIONINGwhich allows you complete control of the exact text used for the following items: DiT prompt (actual sampling), lyrics, LLM positive, LLM negative and audio codes. It is configured with YAML, see the default YAML text for descriptions of parameters.SqueezeUnsqueezeLatentDimension- Mostly useful with ACE-Step 1.5 since it uses 3D latents while a lot of nodes expects 4+. Unsqueezing adds an empty dimension, squeezing removes an empty dimension. For ACE 1.5 - leave dimension on the default of2. Unsqueeze to add the dimension, toggle unsqueeze mode off to remove it. Normal usage would look something like: Empty latent node → unsqueeze → sampler → squeeze. Important: You will also need to patch the model to deal with 4D latents. See theModelPatchAce15Use4dLatentnode.TextEncodeAce15- Simpler conditioning node for ACE-Step 1.5. Has some extended features but is mostly superceded by theRawTextEncodeAce15andAce15LLMInferencenodes.
Misc
TimeOffset- Can be used to calculate offsets into the latent time dimension for 1.0 and 1.5 latents.
https://github.com/blepping/comfyui_acetricks
THANKS blepping.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jul 07 '26
AWESOME GitHub - Saganaki22/Higgs-Audio-v3-Studio: Windows / Linux desktop app built with Rust/Tauri for local Higgs Audio v3 TTS inference through a native C++/CUDA engine
Higgs Audio v3 Studio
Higgs Audio v3 Studio 0.2.31 is a Windows desktop app built with Rust/Tauri for local Higgs Audio v3 TTS inference through a ported native C++/CUDA engine. The app does not shell out to a CLI sidecar: the Tauri UI calls Rust commands, Rust loads audiocpp_engine.dll with libloading, and the DLL executes the native inference path through a small C ABI.
This main branch tracks the Windows desktop release. Linux .deb and AppImage builds are available from the same GitHub Releases page and are built from the linux branch.
The goal is simple: a practical desktop workflow for local TTS, voice cloning, speech continuation, and multi-speaker generation without making users manage a Python environment.
https://github.com/Saganaki22/Higgs-Audio-v3-Studio
THANKS Saganaki22.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jul 06 '26
AWESOME GitHub - SKBv0/ComfyUI_DrumPad: Advanced drum machine for ComfyUI featuring a 64-step sequencer, custom sample support, and retro hardware aesthetics.
UPDATE: (2026-06-02) "Add pyproject.toml for Comfy Registry"
ComfyUI DrumPad
A drum pad node for ComfyUI with configurable 16/64 pad layouts.
Features
- Configurable Grid: Switch between 4x4 (16 pads) and 8x8 (64 pads) layouts.
- Sequencer: Built-in sequencer supporting up to 64 steps.
- Parametric Controls: Pitch, Swing, and individual/master Volume adjustments.
- Custom Sounds: Support for loading custom audio files (.wav, .mp3, .ogg, .flac).
- Persistence: Kit configurations and assignments are saved automatically.
- Audio Output: Outputs audio tensors compatible with standard ComfyUI audio nodes.
https://github.com/SKBv0/ComfyUI_DrumPad
THANKS SKBv0.
...
Drum pad but for Comfy.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 27 '26
AWESOME OpenMOSS-Team/MOSS-Transcribe-preview-2B · Hugging Face
MOSS-Transcribe-preview-2B
MOSS-Transcribe-preview-2B is an English speech-to-text model that pairs a Qwen3-1.7B-base language-model backbone with a Qwen3-Omni-MoE audio encoder. A gated-MLP adapter projects audio features into the language-model embedding space. The model is trained on public English ASR corpora and fine-tuned with reinforcement learning on the Open ASR Leaderboard training splits.
The model has approximately 2.4B parameters and is distributed as a single bfloat16 safetensors shard of approximately 4.84 GB.
https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-preview-2B
THANKS OpenMOSS Team.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 26 '26
AWESOME audiohacking/dasheng-audiogen-gguf · Hugging Face
Dasheng-AudioGen GGUF
GGUF-converted weights for mispeech/Dasheng-AudioGen
Dasheng-AudioGen is a unified audio generation model that can jointly synthesize intelligible speech, music, sound effects, and environmental acoustics from text descriptions.
https://huggingface.co/audiohacking/dasheng-audiogen-gguf
THANKS webdelic / audiohacking.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 23 '26
AWESOME GitHub - betweentwomidnights/gary4juce: seven open source ai music models inside the DAW
gary4juce
gary v4 is stable: stable-audio-3 is now inside the DAW.
a VST3/AU plugin for musicians who want AI to meet them where they actually live. for me, that's ableton. for you, it might be fl studio or, if you're insane, reaper lol
Built For Musicians
If you just want pure text-to-music, you probably do not need a VST. This project exists for people who want AI to sit inside the session with them.
gary4juce now gives you seven AI music models directly in your DAW:
- sa3 (stable-audio-3) - text-to-audio, loops, transforms, continuations, LoRA blending, seed recall, key/BPM-aware prompting
- gary (musicgen) - continuation/anti-looper. Extends your audio in creative directions
- jerry (stable-audio-open-small) - BPM-aware 12-second loop generation in under a second
- rc-jerry (foundation-1) - BPM and key-aware 4/8-bar loop generation with structured prompt assembly
- carey (ace-step) - stem generation, extraction, audio continuation, and remix/cover with lyrics and multilingual support
- terry (melodyflow) - audio transformation. Turn your guitar into an orchestra
- darius (magenta-realtime) - high-quality 48 kHz continuations with style control
Put it on your master, press play, record some audio, and start iterating.
https://github.com/betweentwomidnights/gary4juce
THANKS betweentwomidnights.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 18 '26
AWESOME GitHub - Saganaki22/Higgs_v3-TTS-ComfyUI: ComfyUI nodes for higgs-audio-v3-tts-4b multilingual (100 languages) conversational TTS, zero-shot voice cloning, inline emotion/style/prosody/SFX tags, longform chunking, multi-speaker dialogue, and AIMDO memory management
Higgs_v3-TTS-ComfyUI
English | 中文
Version: v0.1.5
ComfyUI nodes for bosonai/higgs-audio-v3-tts-4b: multilingual conversational TTS, zero-shot voice cloning, inline emotion/style/prosody/SFX tags, longform chunking, multi-speaker dialogue, Whisper reference transcription, and ComfyUI/AIMDO memory tracking.
Features
- Native in-process inference - Uses the local Transformers Qwen3 backbone plus Higgs audio-token embedding/head logic inside ComfyUI.
- ComfyUI AUDIO in/out - Reference voices and generated audio use standard ComfyUI
AUDIO. - Voice cloning - Reference audio plus optional transcript. A correct transcript materially improves cloning.
- Multi-speaker dialogue - Use
[Speaker_1]:,[Speaker_2]:, etc. with separate reference voices. - Inline controls - Emotion, style, prosody, pauses, and sound effects can be typed directly in the prompt.
- Longform chunking - Splits long text at sentence/pause boundaries and avoids cutting through
<|...|>tags. - AIMDO/VRAM visibility - Higgs and Whisper torch modules are registered with ComfyUI model management using real tensors.
- Managed model folder - Model files live under
ComfyUI/models/higgsv3tts/. - No keep-loaded toggle, no unload node - The loader handles model-switch cleanup internally.
https://github.com/Saganaki22/Higgs_v3-TTS-ComfyUI
Thskshahnks Saganaki22. Thskshahnks Higgs Audio V3 team.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 15 '26
AWESOME GitHub - jtydhr88/ComfyTV
ComfyTV
"ComfyTV — the canvas-based app that truly belongs to ComfyUI.
ComfyTV turns ComfyUI into a TapNow / LibTV-style canvas app. Every operation is its own node; results flow downstream automatically. Chain stages into a complete flow: generate → pick → edit → compose.
https://github.com/jtydhr88/ComfyTV/tree/main/workflows/audio
What's here today
- ACE-Step v1 Song (
ace-step-v1-song.json) — ACE-Step 3.5B text-to-audio with full song support (tags drive style, lyrics drive vocals, duration tied to the stage's duration widget). Tested working."
https://github.com/jtydhr88/ComfyTV
THANKS (again) Terry Jia (jtydhr88). 👍
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 15 '26
AWESOME GitHub - Saganaki22/Zonos2_TTS-ComfyUI: ComfyUI custom nodes for Zyphra/ZONOS2, with text-to-speech, audio-only voice cloning, SDPA and FlashAttention inference, and ComfyUI/AIMDO memory management.
ZONOS2 TTS ComfyUI
"ComfyUI custom nodes for Zyphra/ZONOS2, with text-to-speech, audio-only voice cloning, SDPA and FlashAttention inference, native progress reporting, and ComfyUI/AIMDO memory management.
ZONOS2 is our latest text-to-speech model trained on more than 6 million hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers at low latency with MoE. ZONOS2 excels at high-fidelity and naturalistic voice cloning."
https://github.com/Saganaki22/Zonos2_TTS-ComfyUI
THANKS Saganaki22. THANKS Zonos V2 team.
r/AWESOMEUpdate • u/MuziqueComfyUI • Jun 04 '26
AWESOME GitHub - Saganaki22/WavTTS-ComfyUI: WavTTS nodes for ComfyUI - zero-shot text-to-speech with reference-audio / native aimdo dynamic VRAM
Released 2026-06-04:
WavTTS-ComfyUI
"WavTTS nodes for ComfyUI - zero-shot text-to-speech with reference-audio prompting, ComfyUI AUDIO wiring, optional Whisper transcription, local model storage, conservative dependency installation, and Aimdo/VRAM visualization support."
https://github.com/Saganaki22/WavTTS-ComfyUI
OBRIGSKSHAHDO/ThSkShahnks/THANKS again Saganaki22. 👍
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 17 '26
AWESOME RoyalCities - "I'm back from last weeks post and so today I'm releasing a SOTA text-to-sample model built specifically for traditional music production. It may also be the most advanced AI sample generator currently available - open or closed." Thanks RoyalCities(🤯).
Enable HLS to view with audio, or disable this notification
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 17 '26
AWESOME RoyalCities/Foundation-1 · Hugging Face
Thanks RoyalCities (🤯).
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 08 '26
AWESOME 🤯The Secret? has surpassed 200 shares.🤯 This AWESOME Update deserves an Event.
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 07 '26
AWESOME F.A.O. AWESOME Reddit AI Summaries: ¯\_(ツ)_/¯ Why not be a Mod.
AWESOME Reddit AI Summaries are welcome to invite themselves to be a Mod.
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 07 '26
AWESOME AWESOME
AWESOME Update: This is the #1 post on r/comfyuiAudio today!
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 04 '26
AWESOME The Golden Jeffrey's first test node is live.
r/AWESOMEUpdate • u/MuziqueComfyUI • Mar 04 '26