UniRouteUniRoute

omnihuman-1-5

Commercial use

OmniHuman 1.5 is ByteDance's AI avatar and digital human model that transforms a single image and audio into realistic talking videos with natural lip sync, facial expressions, and lifelike motion.

Model Type:
Pricing: from $0.189/s
View all pricing

omnihuman-1-5/human-identification USD / image

Option / channelOfficialCloud vendor

omnihuman-1-5/subject-detection USD / image

Option / channelOfficialCloud vendor

omnihuman-1-5 USD / second

Option / channelOfficialCloud vendor
Standard$0.27$0.189

Input

image_url
File 1
image_url 1

File uploads require configured image storage. A public HTTPS image URL can be used directly.

Portrait image URL, supports any aspect ratio with subjects including people/pets/anime, etc.
0
To have a specific subject in the image speak, use 'Subject Detection' to get the corresponding mask image and pass it as input.0
audio_url
File 1

File uploads require configured image storage. A public HTTPS image URL can be used directly.

Audio URL. Duration must be < 60 seconds (recommended ≤15 seconds; exceeding this will cause degradation).
Prompt text, limited to Chinese/English/Japanese/Korean/Spanish/Indonesian, recommended ≤300 characters.65/300
Output video resolution, default 1080.

Fast mode, sacrifices some quality to speed up generation.

Random seed. Default is -1 (random). When using the same positive integer and keeping all other parameters identical, the result will be highly consistent (with very high probability).

Output

output typevideoExample output

README

Complete guide to using omnihuman-1-5