On-prem models
Voice training and fine-tuning
Fine-tunes a dedicated voice from your recordings on a local GPU. A trained voice turns text into speech, with controls for stability, speed and emotion; a pronunciation table corrects Taiwanese Mandarin readings.
What this system does
- Each voice gets its own adapter; every epoch is kept, so you can play the same line on each version and pick one
- Zero-shot cloning for when there is not enough audio yet: 3–15 seconds of clean speech gives a first preview
- Pronunciation table for Taiwanese Mandarin: whole-word replacements and polyphone overrides, applied on save
- HTTP API and MCP tools; AgentHub synthesises its LINE voice broadcasts with it
Inside the product
Live screens. Fields containing customer data are masked. Tap a screenshot to see it full size.

Trained-voice synthesis: pick a trained voice, type the text, add pauses or an emotion, and synthesise. 
Instant ICL cloning: upload 3–15 seconds of reference audio and its transcript to hear the voice without training. 
Pronunciation fixes: a whole-word replacement table for Taiwanese usage, applied to every voice on save.
Virtual characters speaking
What a trained voice looks like on a virtual character: each line is spoken in the character’s own voice, the lips are synced to it, and subtitles are burned in.
Greeting: an 8-second self-introduction in a record shop, asking viewers to follow; the subtitles keep pace with the voice. Lip-sync close-up: a man talks straight to the camera, and even this close the mouth matches the voice. Weekend vlog: three small ways to slow down, told while walking through a park — 37 seconds across several shots, with the same face and voice throughout.
More work
Cases from the same service come first.

On-prem models
Virtual human and AI media generation platform
A generation platform on our own GPU server. Each virtual human has its own LoRA, so the face stays the same across scenes and outfits; they shoot everyday clips and act in multi-character sketches, lip-synced in their own voices. It also turns a topic, a news article or a slide deck into a narrated, subtitled video, and can produce a whole instrumental album. Images, video, voice and music are all generated on the local GPU.
- Virtual humans: one LoRA per character, the same face across scenes and outfits, each with its own voice and persona
- Clips and sketches: AI writes the script, plans the storyboard, generates frames and shoots the video, and you can stop to review, edit or redo any step
ComfyUI · Z-Image · Krea 2 · Wan 2.2 · LTX · ACE-Step · local GPU
See insideOn-prem AI model solutions
E-commerce
eSIM 226 storefront
A travel eSIM store in eight languages with multi-currency pricing and two Taiwanese payment gateways. Paid orders are provisioned with the upstream supplier automatically and the install email goes out on its own.
- Eight locales and multiple display currencies, so Japanese buyers see yen
- Automatic provisioning across two suppliers, with failover
Next.js · PostgreSQL · Keycloak · ECPay/NewebPay · Kubernetes

Admin system
eSIM back office
The operations side of the same store: products, pricing, orders, refunds, suppliers and payments in one place. It actively surfaces problems like "paid but never fulfilled" instead of waiting to be asked.
- Orders, revenue, unpaid and unfulfilled counts at a glance
- AI pricing suggestions and bulk listing changes
Next.js · PostgreSQL · Keycloak · GA4 Data API
See insideE-commerce site