Skip to content

On-prem AI model solutions

Virtual humans, AI video, voice and music, generated on your own GPUs

We host open models on your GPU server: each virtual human gets a dedicated LoRA so the face stays the same across scenes and outfits, and shoots everyday clips or acts in short sketches lip-synced in its own voice. A topic, a news article or a slide deck becomes a narrated, subtitled video, and you can also train a custom voice, generate background music and transcribe recordings. Images, video and audio are generated on your machines, with no per-call billing.

Generated clips

Clips this platform actually generated: first an AI ad; then three virtual characters speaking, each line voiced in the character’s own voice, lip-synced and subtitled; and last, three image-to-video clips (MiniMax H3) starting from a single keyframe of a virtual human, 5 seconds each, unedited.

  • AI ad: a 58-second family-travel ad about taking a three-year-old to Kansai — 11 scenes, each with a title and subtitles, ending on “request the itinerary”.
  • Greeting: an 8-second self-introduction in a record shop, asking viewers to follow; the subtitles keep pace with the voice.
  • Lip-sync close-up: a man talks straight to the camera, and even this close the mouth matches the voice.
  • Weekend vlog: three small ways to slow down, told while walking through a park — 37 seconds across several shots, with the same face and voice throughout.
  • Park path: she walks toward the camera smiling; her face, clothes and the light stay consistent with the keyframe it started from.
  • Packing: one hand on the suitcase, she laughs, looks down and back up at the camera; the room and morning light hold steady throughout.
  • Café: she lifts the cup and sips while the camera slowly pushes in to her smile; the music and ambient sound are generated by the model too.

A good fit when

  • You produce a lot of images, audio or video and per-call cloud pricing no longer adds up
  • Recordings, faces or transcripts are sensitive and cannot go to a third party
  • You need one consistent brand voice or presenter across every output

Not a fit when

  • You produce a handful a month, where an off-the-shelf cloud service is cheaper
  • You do not hold the rights to the voice or likeness

What you get

  1. 01Virtual humans: a dedicated character LoRA trained on licensed material, so the face stays the same across scenes and outfits, with a voice and persona set up
  2. 02Virtual human videos: vertical everyday clips, or short sketches with 2–4 characters on screen; lines are lip-synced in each character's voice, and you can stop to revise the script, storyboard and frames at every step
  3. 03AI production: start from a topic, news article, blog post or PDF/slides; AI writes the script, picks the shots, adds narration and subtitles, and you finish the cut on a timeline
  4. 04Generation workbench: text-to-image, image-to-image, retouching, outfit swaps, image-to-video, music and voice, with switchable models; results go straight into the media library
  5. 05Voice and editing: custom voices trained from your recordings; meetings and support calls transcribed with timestamps and speaker separation; podcasts and videos trimmed of silences and repeats, exported as short versions with subtitles
  6. 06Deployment and interfaces: runs on your GPU host or a cloud GPU we buy for you, with a GPU queue, VRAM monitoring, an API and MCP tools; scripts can be written by Claude or entirely by local models

What it looks like

Screens from Virtual human and AI media generation platform. Click an image to see it full size.

  • Virtual humans: each character has a dedicated LoRA, a persona and a social handle, with asset completeness at a glance.
    Virtual humans: each character has a dedicated LoRA, a persona and a social handle, with asset completeness at a glance.
  • Everyday videos: pick a virtual human, say what to film today, and get a vertical clip that looks casually shot.
    Everyday videos: pick a virtual human, say what to film today, and get a vertical clip that looks casually shot.
  • Short sketches: pick 2–4 actors and one idea; the script, storyboard, frames and video are confirmed step by step.
    Short sketches: pick 2–4 actors and one idea; the script, storyboard, frames and video are confirmed step by step.
  • Generation workbench: tabs for images, video, music and voice, switchable image models, results saved to the media library.
    Generation workbench: tabs for images, video, music and voice, switchable image models, results saved to the media library.
  • AI production: start from a topic, news story or blog post; several AI agents write the script, then narration and subtitles finish the video.
    AI production: start from a topic, news story or blog post; several AI agents write the script, then narration and subtitles finish the video.
  • Relaxation albums: four AI roles arrange a track list that moves from tense to calm, and ACE-Step generates each track.
    Relaxation albums: four AI roles arrange a track list that moves from tense to calm, and ACE-Step generates each track.
More about Virtual human and AI media generation platform

Related articles

Write-ups on our founder's blog, covering how it was done and the problems along the way.

Related work

Systems we run ourselves or have published that are the same kind of thing as this service.

  • Voice training and fine-tuning

    On-prem models

    Voice training and fine-tuning

    Fine-tunes a dedicated voice from your recordings on a local GPU. A trained voice turns text into speech, with controls for stability, speed and emotion; a pronunciation table corrects Taiwanese Mandarin readings.

    • Each voice gets its own adapter; every epoch is kept, so you can play the same line on each version and pick one
    • Zero-shot cloning for when there is not enough audio yet: 3–15 seconds of clean speech gives a first preview

    Qwen3-TTS · CosyVoice3 · LoRA · local GPU · MCP

    See insideOn-prem AI model solutionsVisit site
  • Seal-TTS Taiwanese voice platform

    Seal-TTS Taiwanese voice platform

    A voice platform on a local GPU: train dedicated voices from recordings, zero-shot cloning, multi-speaker podcasts, speech recognition, vocal separation and pronunciation fixes. AgentHub’s LINE voice broadcasts are synthesised with it.

    • Qwen3-TTS
    • CosyVoice3
    • Local GPU
    • MCP
  • AI Podcast Cut

    Removes filler words and stumbles from a podcast: local faster-whisper recognition, demucs vocal separation, level balancing that keeps breaths, and each seam can be auditioned before export. Everything runs locally.

    • faster-whisper
    • demucs
    • Rust
    • ffmpeg
    • Open source
  • AI Video Cut

    Pick an object in a video by typing, clicking or letting AI choose; it is tracked frame by frame, then blurred, replaced with your own image or video, or removed. Also does timeline edits and local subtitles, all on your own GPU.

    • SAM 2.1
    • faster-whisper
    • MCP
    • Local GPU
    • MIT
See all work

Frequently asked

It depends on the model and the quality you want. A few clean minutes usually gets a first version; half an hour to an hour makes it steadier. We build a sample from whatever you already have so you can listen before committing.