On-prem AI model solutions
Virtual humans, AI video, voice and music, generated on your own GPUs
We host open models on your GPU server: each virtual human gets a dedicated LoRA so the face stays the same across scenes and outfits, and shoots everyday clips or acts in short sketches lip-synced in its own voice. A topic, a news article or a slide deck becomes a narrated, subtitled video, and you can also train a custom voice, generate background music and transcribe recordings. Images, video and audio are generated on your machines, with no per-call billing.
Generated clips
Clips this platform actually generated: first an AI ad; then three virtual characters speaking, each line voiced in the character’s own voice, lip-synced and subtitled; and last, three image-to-video clips (MiniMax H3) starting from a single keyframe of a virtual human, 5 seconds each, unedited.
AI ad: a 58-second family-travel ad about taking a three-year-old to Kansai — 11 scenes, each with a title and subtitles, ending on “request the itinerary”. Greeting: an 8-second self-introduction in a record shop, asking viewers to follow; the subtitles keep pace with the voice. Lip-sync close-up: a man talks straight to the camera, and even this close the mouth matches the voice. Weekend vlog: three small ways to slow down, told while walking through a park — 37 seconds across several shots, with the same face and voice throughout. Park path: she walks toward the camera smiling; her face, clothes and the light stay consistent with the keyframe it started from. Packing: one hand on the suitcase, she laughs, looks down and back up at the camera; the room and morning light hold steady throughout. Café: she lifts the cup and sips while the camera slowly pushes in to her smile; the music and ambient sound are generated by the model too.
A good fit when
- You produce a lot of images, audio or video and per-call cloud pricing no longer adds up
- Recordings, faces or transcripts are sensitive and cannot go to a third party
- You need one consistent brand voice or presenter across every output
Not a fit when
- You produce a handful a month, where an off-the-shelf cloud service is cheaper
- You do not hold the rights to the voice or likeness
What you get
- 01Virtual humans: a dedicated character LoRA trained on licensed material, so the face stays the same across scenes and outfits, with a voice and persona set up
- 02Virtual human videos: vertical everyday clips, or short sketches with 2–4 characters on screen; lines are lip-synced in each character's voice, and you can stop to revise the script, storyboard and frames at every step
- 03AI production: start from a topic, news article, blog post or PDF/slides; AI writes the script, picks the shots, adds narration and subtitles, and you finish the cut on a timeline
- 04Generation workbench: text-to-image, image-to-image, retouching, outfit swaps, image-to-video, music and voice, with switchable models; results go straight into the media library
- 05Voice and editing: custom voices trained from your recordings; meetings and support calls transcribed with timestamps and speaker separation; podcasts and videos trimmed of silences and repeats, exported as short versions with subtitles
- 06Deployment and interfaces: runs on your GPU host or a cloud GPU we buy for you, with a GPU queue, VRAM monitoring, an API and MCP tools; scripts can be written by Claude or entirely by local models
What it looks like
Screens from Virtual human and AI media generation platform. Click an image to see it full size.

Virtual humans: each character has a dedicated LoRA, a persona and a social handle, with asset completeness at a glance. 
Everyday videos: pick a virtual human, say what to film today, and get a vertical clip that looks casually shot. 
Short sketches: pick 2–4 actors and one idea; the script, storyboard, frames and video are confirmed step by step. 
Generation workbench: tabs for images, video, music and voice, switchable image models, results saved to the media library. 
AI production: start from a topic, news story or blog post; several AI agents write the script, then narration and subtitles finish the video. 
Relaxation albums: four AI roles arrange a track list that moves from tense to calm, and ACE-Step generates each track.
Related articles
Write-ups on our founder's blog, covering how it was done and the problems along the way.
Related work
Systems we run ourselves or have published that are the same kind of thing as this service.

On-prem models
Voice training and fine-tuning
Fine-tunes a dedicated voice from your recordings on a local GPU. A trained voice turns text into speech, with controls for stability, speed and emotion; a pronunciation table corrects Taiwanese Mandarin readings.
- Each voice gets its own adapter; every epoch is kept, so you can play the same line on each version and pick one
- Zero-shot cloning for when there is not enough audio yet: 3–15 seconds of clean speech gives a first preview
Qwen3-TTS · CosyVoice3 · LoRA · local GPU · MCP

Seal-TTS Taiwanese voice platform
A voice platform on a local GPU: train dedicated voices from recordings, zero-shot cloning, multi-speaker podcasts, speech recognition, vocal separation and pronunciation fixes. AgentHub’s LINE voice broadcasts are synthesised with it.
- Qwen3-TTS
- CosyVoice3
- Local GPU
- MCP

Removes filler words and stumbles from a podcast: local faster-whisper recognition, demucs vocal separation, level balancing that keeps breaths, and each seam can be auditioned before export. Everything runs locally.
- faster-whisper
- demucs
- Rust
- ffmpeg
- Open source

Pick an object in a video by typing, clicking or letting AI choose; it is tracked frame by frame, then blurred, replaced with your own image or video, or removed. Also does timeline edits and local subtitles, all on your own GPU.
- SAM 2.1
- faster-whisper
- MCP
- Local GPU
- MIT