Skip to content

Backend reference

Backends

Recipe Name Selectable backend Uses ctx_size Backends
acestep ACE-Step yes no cuda, rocm, vulkan
ds4 DwarfStar4 (experimental) no yes rocm
flm FastFlowLM NPU no yes npu
kokoro Kokoro no no cpu, metal
llamacpp Llama.cpp GPU yes yes cpu, cuda, metal, rocm, system, vulkan
llamacpp-hrx HRX GPU (experimental) no yes hrx
moonshine Moonshine no no cpu
onnxruntime ONNX Runtime no no cpu
openmoss OpenMOSS TTS yes no cuda, vulkan
ryzenai-llm Ryzen AI LLM no yes npu
sd-cpp StableDiffusion.cpp yes no cpu, cuda, metal, rocm, vulkan
thenoise TheNoise ROCm yes no rocm
thinksound ThinkSound yes no cuda, rocm, vulkan
trellis TRELLIS.2 yes no cuda, rocm, vulkan
vllm vLLM ROCm (experimental) yes yes rocm
whispercpp Whisper.cpp yes no cpu, metal, npu, rocm, vulkan

Support matrix

Recipe Backend OS Device families
acestep cuda linux, windows nvidia_gpu
acestep vulkan linux, windows amd_gpu; cpu (x86_64); nvidia_gpu
acestep rocm linux, windows amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X)
ds4 rocm linux amd_gpu (gfx1151)
flm npu linux, windows amd_npu (XDNA2)
kokoro metal macos metal
kokoro cpu linux, windows cpu (x86_64)
llamacpp system linux cpu (arm64, x86_64)
llamacpp metal macos metal
llamacpp cuda linux, windows nvidia_gpu (sm_100, sm_120, sm_121, sm_75, sm_80, sm_86, sm_89, sm_90)
llamacpp vulkan linux, windows amd_gpu; cpu (arm64, x86_64)
llamacpp rocm linux, windows amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X, gfx908, gfx90a, gfx942, gfx950)
llamacpp cpu linux, windows cpu (arm64, x86_64)
llamacpp-hrx hrx linux amd_gpu (gfx1100, gfx1151)
moonshine cpu windows cpu (x86_64)
moonshine cpu linux cpu (arm64, x86_64)
moonshine cpu macos cpu (arm64)
onnxruntime cpu windows cpu (x86_64)
onnxruntime cpu linux cpu (arm64, x86_64)
onnxruntime cpu macos cpu (arm64)
openmoss cuda linux, windows nvidia_gpu
openmoss vulkan linux, windows amd_gpu; cpu (x86_64); nvidia_gpu
ryzenai-llm npu windows amd_npu (XDNA2)
sd-cpp metal macos metal
sd-cpp cuda linux, windows nvidia_gpu (sm_100, sm_120, sm_121, sm_75, sm_80, sm_86, sm_89, sm_90)
sd-cpp vulkan linux, windows amd_gpu; cpu (x86_64); nvidia_gpu
sd-cpp rocm linux amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X)
sd-cpp cpu linux, windows cpu (x86_64)
thenoise rocm linux amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X)
thinksound cuda linux, windows nvidia_gpu
thinksound vulkan linux, windows amd_gpu; cpu (x86_64); nvidia_gpu
thinksound rocm linux, windows amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X)
trellis cuda linux, windows nvidia_gpu
trellis vulkan linux, windows amd_gpu; cpu (x86_64); nvidia_gpu
trellis rocm linux, windows amd_gpu (gfx103X, gfx110X, gfx1150, gfx1151, gfx1152, gfx120X)
vllm rocm linux amd_gpu (gfx110X, gfx1150, gfx1151, gfx120X)
whispercpp npu windows amd_npu (XDNA2)
whispercpp metal macos metal
whispercpp vulkan linux, windows amd_gpu; cpu (x86_64)
whispercpp rocm linux, windows amd_gpu (gfx110X, gfx1150, gfx1151, gfx120X)
whispercpp cpu linux, windows cpu (x86_64)

Note: The llamacpp rocm row lists linux, windows for the family as a whole, but MI350X (gfx950) is currently gated to Linux + stable channel only — the Windows TheRock distribution and the ROCm nightly build for gfx950 are not yet published, so gfx950 installs are rejected on Windows and on the nightly channel. The OS column reflects the row's overall reach; the per-architecture restriction is enforced by the backend's install gate.

Recipe options

acestep — ACE-Step

Option CLI flag Type Default Description
acestep_backend --acestep BACKEND "" ACE-Step backend to use

ds4 — DwarfStar4 (experimental)

Option CLI flag Type Default Description
ctx_size --ctx-size SIZE -1 Context size for the model
ds4_args --ds4-args ARGS "" Custom arguments to pass to ds4-server

flm — FastFlowLM NPU

Option CLI flag Type Default Description
ctx_size --ctx-size SIZE -1 Context size for the model
flm_args --flm-args ARGS "" Safe flm serve tuning args: --pmode, --prefill-chunk-len, --img-pre-resize, --socket, --q-len, --preemption

llamacpp — Llama.cpp GPU

Option CLI flag Type Default Description
ctx_size --ctx-size SIZE -1 Context size for the model
llamacpp_backend --llamacpp BACKEND "" LlamaCpp backend to use
llamacpp_device --llamacpp-device DEVICES "" Comma-separated list of accelerator devices to use (e.g. Vulkan0)
llamacpp_args --llamacpp-args ARGS "" Custom arguments to pass to llama-server

llamacpp-hrx — HRX GPU (experimental)

Option CLI flag Type Default Description
ctx_size --ctx-size SIZE -1 Context size for the model
hrx_args --hrx-args ARGS "" Custom arguments to pass to the HRX llama-server

moonshine — Moonshine

Option CLI flag Type Default Description
moonshine_args --moonshine-args ARGS "" Custom arguments to pass to moonshine-server

onnxruntime — ONNX Runtime

Option CLI flag Type Default Description
onnxruntime_args --onnxruntime-args ARGS "" Custom arguments to pass to ort-server

openmoss — OpenMOSS TTS

Option CLI flag Type Default Description
openmoss_backend --openmoss BACKEND "" OpenMOSS TTS backend to use

sd-cpp — StableDiffusion.cpp

Option CLI flag Type Default Description
sd-cpp_backend --sdcpp BACKEND "" SD.cpp backend to use
sdcpp_args --sdcpp-args ARGS "" Custom arguments to pass to sd-server (must not conflict with managed args)
steps — SIZE 20 Number of diffusion steps
cfg_scale — SIZE 7.0 Classifier-free guidance scale
width — SIZE 512 Output image width
height — SIZE 512 Output image height
sampling_method — ARGS "" Sampling method
flow_shift — SIZE 0.0 Flow shift

thenoise — TheNoise ROCm

Option CLI flag Type Default Description
thenoise_backend --thenoise BACKEND "" TheNoise backend to use
steps — SIZE 20 Number of denoising steps
cfg_scale — SIZE 7.0 CFG scale (<= 1.0 disables CFG)
width — SIZE 512 Output image width
height — SIZE 512 Output image height
sampler — ARGS "" Denoising solver (euler | er_sde)
negative_prompt — ARGS "" Negative prompt
qwen_vae_enhance — BOOL false Nyquist notch post-filter (removes 2px grid artifacts)
film_grain — SIZE 0.0 Film grain strength (0.0-10.0)
sharpening — SIZE 0.0 RCAS sharpening strength (0.0-1.0)
lora_specs — ARGS "" Comma-separated LoRA specs, e.g. "style:0.8,sub/detail:0.5"

thinksound — ThinkSound

Option CLI flag Type Default Description
thinksound_backend --thinksound BACKEND "" ThinkSound backend to use

trellis — TRELLIS.2

Option CLI flag Type Default Description
trellis_backend --trellis BACKEND "" Trellis backend to use
trellis_args --trellis-args ARGS "" Custom arguments to pass to trellis-server

vllm — vLLM ROCm (experimental)

Option CLI flag Type Default Description
ctx_size --ctx-size SIZE -1 Context size for the model
vllm_backend --vllm BACKEND "" vLLM backend to use
vllm_args --vllm-args ARGS "" Custom arguments to pass to vllm-server

whispercpp — Whisper.cpp

Option CLI flag Type Default Description
whispercpp_backend --whispercpp BACKEND "" WhisperCpp backend to use
whispercpp_args --whispercpp-args ARGS "" Custom arguments to pass to whisper-server

Implementation notes

ACE-Step (acestep)

ace-server exposes an asynchronous job API: POST /lm or POST /synth returns a job id immediately, GET /job?id=N polls the status, and GET /job?id=N&result=1 fetches the finished result. AceStepServer::run_job wraps this submit/poll/fetch cycle with a ceiling of roughly 20 minutes per stage at a 1-second poll cadence. Synth results arrive as multipart/mixed (an audio part plus a latent part); Lemonade extracts the first audio part.

Vocals are a two-stage pipeline. POST /lm with lm_mode: "generate" runs the ACE-Step language model, which turns the caption and lyrics into audio codes plus LM-filled metadata, returned as a JSON array of enriched requests. That array is accepted by POST /synth verbatim, so Lemonade feeds it through unchanged. POST /synth on its own is the DiT-only instrumental path — it has no language model and cannot sing, so a lyrics value (other than the sentinel below) is what routes a request through /lm first.

[Instrumental] — any case, surrounding whitespace ignored — is ACE-Step's sentinel for the no-vocals path, matching the Python reference implementation. Instrumental requests send the sentinel explicitly rather than an empty string because the synth stage also feeds the lyrics text into its conditioning.

Errors from audio_generations are written into the response sink as a JSON error payload; the endpoint handler turns that into an HTTP error instead of shipping it as audio.

The model download fetches the DiT checkpoint variant plus three companions when present in the repo: the language model (acestep-5Hz-lm-4B-Q8_0.gguf, required for vocals and auto-lyrics), the Qwen3 text encoder, and the VAE. The checkpoint path handed to --models is the directory of GGUFs; ace-server scans it by architecture, and --keep-loaded keeps models resident across requests.

OpenMOSS (openmoss)

moss-tts-server hosts exactly one --model per process, so the voice generator that ships as a component of the speech model cannot be served by the same process. OpenMossServer runs the cascade one model at a time: the speech process is stopped, a transient process is spawned on the voice-generator checkpoint to render a single reference sample, and the speech process is brought back. Holding both would require a card that fits the pair, a strictly harder requirement than running the model that was actually asked for. For the same reason, server_models.json keeps size as the peak resident requirement of the main speech model rather than adding the non-resident VoiceGen download size. Rendered samples are cached per description for the life of the load, so repeating a description costs nothing. The speech process is restarted even when design fails, so an unsuccessful design cannot leave a loaded model with no process behind it.

The legacy MOSS-VoiceGen registry entry remains as a compatibility model for existing clients; integrated voice design on OpenMOSS-TTS and MOSS-TTS-Local does not depend on that standalone entry.

request_mutex_ serialises a process-changing voice-design request against normal OpenMOSS inference, load(), and unload(). Speech and SFX hold it shared; voice design and lifecycle changes hold it exclusively. During the intentional speech -> VoiceGenerator -> speech swap, process_swap_in_progress_ keeps is_backend_alive() true so the router does not mistake the temporarily empty speech-process handle for a crash. A request arriving in that window is accepted and then waits on request_mutex_ until speech is restored. start_speech_process() therefore calls stop_speech_process() rather than unload() on a failed readiness wait, because unload() takes the same exclusive lock.

The transient design process is polled for readiness through its own /health rather than the shared wait_for_ready(), since it does not own port_; the poll also checks ProcessManager::is_running() each round and reports the child's exit code, so a subprocess that dies during model loading returns that error instead of polling to the timeout. spawn() uses find_free_port() instead of choose_port() for the same reason — choose_port() assigns port_, and the caller decides which process port_ addresses.

Reference-conditioned TTS can exceed OpenMOSS's default 8192-token context. Lemonade intentionally keeps the upstream context default instead of forcing a larger allocation, preserving compatibility with memory-constrained GPUs. OpenMOSS v0.3 chunks prefill batches, but a complete prompt still has to fit n_ctx; unusually long references may therefore be rejected by upstream with guidance to shorten the reference or raise the context size.

Lemonade does not inject max_audio_frames into OpenMOSS speech requests. OpenMOSS v0.3 owns this policy: its VoiceGenerator derives a text-based duration window when needed, while MOSS-TTS/MOSS-TTSD keep their reference-aware/native stop behavior. Keeping that distinction in the backend avoids truncating conditioned or multi-speaker speech.

Voice design is opt-in through the voice_design_description extension and is never inferred from voice, which keeps its OpenAI-compatible meaning and is forwarded as an instruction. A client sending "voice": "default" gets speech rather than a design run for a voice literally named "default". The field is ignored when the request already carries reference_wav_b64. Request fields are read with a type-checking accessor rather than json::value(), which throws on a type mismatch instead of falling back to the default and would turn a client's wrong-typed field into a 500.

MOSS-SoundEffect uses the same recipe but is an audio-generation model: audio_generations() forwards to the backend's /sfx endpoint, accepting duration/cfg as aliases for seconds/cfg_scale.

llama.cpp (llamacpp)

Lemonade launches llama-server with --parallel 1 and leaves the upstream --cache-ram host prompt cache at its default. LlamaCppServer does not override downsize(): slot erase frees no device memory (the KV buffer is allocated once at model load) and discards the slot's KV state without going through the host-cache save path, so soft idle leaves the slot resident and a resumed conversation reuses its cached prefix.

Model downloads

A checkpoint file can be reached twice during a registry download: once because the backend's select_checkpoint_files claimed it alongside the main weight, and again because it is also declared as its own checkpoint role (the OpenMOSS .extras.gguf sidecars are both). The same bytes either way, so download_from_registry collapses duplicates before counting, or the progress total overshoots.

Backend auto-selection

When a recipe's backend is not pinned in config.json (the backend key is absent or "auto"), the default backend reported in system-info — and used by RecipeOptions when resolving *_backend options — is chosen as follows: the first supported backend in RECIPE_DEFS preference order wins, unless a later supported backend is already locally installed (state installed, update_available, or update_required) while the earlier candidates are merely installable. In that case the first installed one wins. This makes explicitly installing a variant (e.g. the Vulkan build of a GPU backend) an effective override: auto-selection uses what is on disk instead of downloading the preference-order winner. An explicit backend value in config.json always takes precedence over both rules. The llamacpp system variant is never auto-selected unless prefer_system is set.