你兴冲冲打开 MiniMax-H3,输入「一个女孩在雨中奔跑,画面唯美」,点了生成。
等了半分钟,视频出来了。镜头乱晃,女孩的脸走形,没有脚步声,没有雨声,背景音乐像从电梯里偷录的。
你关掉页面,心想:就这?
但问题大概率不在模型,在你的提示词。

2026 年 8 月 3 日,MiniMax 正式开源了新一代全模态生成模型 H3。它能同时理解文字、图片、视频和音频,生成最高 2K 分辨率、最长 15 秒、带原生立体声的视频。简单说,你给它一段描述,它帮你拍一段带声音的短片。
但 H3 不是「你说什么它画什么」的画板,它更像一个听指令干活的全能剧组——你得告诉它拍什么镜头、镜头怎么动、谁在说话、现场什么声音、配什么音乐。指令越精确,出片越稳。
这篇教程,就是教你怎么当好这个「导演」。
01 · 先搞清楚:你要拍哪种片子
H3 把视频生成拆成了四种任务类型,对应四种不同的「拍摄场景」。你先得选对类型,提示词的写法才跟着对。
用拍电影来打比方:
T2VA(文生视频)——你只有剧本,没有图。从零开始拍,自由度最高,但也最容易跑偏。就像导演拿到一个故事大纲,要自己定场景、选角、设计镜头。
I2VA(图生视频)——你有一张开场截图。从这张图往后拍,人物、服装、场景得跟图保持一致。就像导演拿到了第一镜的定妆照,后面的戏得接着演。
FL2VA(首尾帧生视频)——你有开头和结尾两张图。中间怎么过渡,你来写。就像导演说「开场是雨中骑车,结尾是撑伞站着」,中间的转身、开伞动作你来编排。
L2VA(尾帧生视频)——你只有结尾截图。得反推一个合理的前情,让动作自然地落到这张图上。就像导演拿到一个碎玻璃的特写,得倒推手是怎么碰到的、杯子怎么掉的。
四种类型的区别,核心就在「你给不给参考图、给几张、给的是开头还是结尾」。选错类型,模型会一脸懵。
| 任务类型 | 参考图 | 你的角色 | 适合场景 |
|---|---|---|---|
| T2VA | 无 | 从零编剧 | 创意发散、概念短片 |
| I2VA | 1 张(首帧) | 接着图往后拍 | 产品展示、角色动画 |
| FL2VA | 2 张(首+尾) | 编排中间过渡 | 变换、状态改变 |
| L2VA | 1 张(尾帧) | 反推前情 | 悬念揭晓、因果倒叙 |
02 · 提示词长什么样:三个核心部门
不管你选哪种任务类型,提示词的主体结构是一样的。H3 的提示词分两部分:第一部分是指令(告诉模型参考图怎么对齐,T2VA 没有这部分),第二部分是三个核心字段。
继续用剧组的比方——三个核心字段就是三个部门:
integrated_multimodal_description——摄影指导的镜头脚本。画面是什么风格、谁在画面里、镜头怎么运动、谁说了什么话、现场有什么声音,全写在这里。这是篇幅最大、最核心的字段。
overall_soundscape——录音师的现场音笔记。雨声、风声、脚步声、衣服摩擦声、呼吸声、笑声……所有非语言的环境音和动作音,用 1-4 句话概括。注意:对白和唱歌不写这里,它们归摄影指导管。
non_diegetic_music——配乐监督的音乐简报。角色听不到、只有观众能听到的背景音乐。用什么乐器、什么节奏、什么动态变化,用 1-3 句话写清楚。别写「悲伤的音乐」这种抽象情绪词,写具体的乐器和节奏。
integrated_multimodal_description: [Shot 1] 画面风格、构图、人物、动作、对白、镜头运动……
overall_soundscape: 环境音、动作音、非人声的概括。
non_diegetic_music: 背景音乐的乐器、节奏、动态变化。
如果是 I2VA、FL2VA 或 L2VA,在最前面还要加一行对齐指令,告诉模型参考图对应视频的哪个时间点。T2VA 没有这行,直接从三个字段开始。
03 · 镜头脚本:integrated_multimodal_description
这是整篇提示词的灵魂。你要在这里把每一个镜头的画面、动作、对白和现场音沿着时间线写清楚。拆成五个零件来学。
零件一:开场定风格和构图
每个视频的第一个镜头,开头要交代整体风格和初始构图。风格词从这些里选:Cinematic(电影感)、live-action(真人实拍)、2D-animated(二维动画)、3D CG(三维渲染)、claymation(黏土动画)、watercolor(水彩)、vintage film(复古胶片)。
[Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise.
翻译一下:「真人实拍,电影感,一个中远景镜头,画面里是一个面包师在日出前拉开街边小面包店的卷帘门。」
一句话把风格、景别、主体、动作、场景全交代了。这就是你要练的语感。
零件二:分镜和切换
一个视频可以有多个镜头。第一个镜头不加时间戳,后续镜头按顺序编号,每个开头写切换时间:
[Shot 2] At 00:03.500, the camera cuts to a close-up of steam rising from the sliced bread...
切换词用 the camera cuts to 或 the shot cuts to。每次切换都应该带来新信息——换了景别、换了场景、换了视角、或者换了时间。如果只是镜头稍微推近一点,那不叫切镜头,叫镜头运动。
用户明确要求时,也可以用叠化(cross-dissolve)、淡入淡出(fade)、擦除(wipe)这些特殊转场。但默认就用硬切。
零件三:镜头运动(类型 + 幅度 + 速度)
这是小白最容易忽略、但对画面影响最大的一个零件。一个完整的镜头运动描述有三个维度:
运动类型(镜头怎么动)+ 幅度(动多少)+ 速度(动多快)。幅度和速度在中等/正常时可以省略。
| 运动类型 | 意思 | 关键词 |
|---|---|---|
| Zoom In / Out | 机身不动,变焦推远拉近 | Zoom In / Pull Out |
| Push In / Pull Out | 机身前进 / 后退 | Push In / Pull Out |
| Pan Left / Right | 原地左右摇镜头 | Pan Left / Pan Right |
| Truck Left / Right | 整体左右平移 | Truck Left / Truck Right |
| Tilt Up / Down | 原地上下摇 | Tilt Up / Tilt Down |
| Arc Shot | 绕主体弧线运动 | Arc Shot |
| Tracking Shot | 跟拍移动主体 | Tracking Shot |
| Static Shot | 完全不动 | Static Shot |
| POV | 第一人称视角 | POV |
幅度用 with small amplitude(小幅)或 with large amplitude(大幅);速度用 at slow speed(慢速)或 at fast speed(快速)。
关键写法:把镜头运动写成一句自然的英文动作,别堆在句尾当标签。
✓ The camera pushes in with small amplitude at slow speed toward the folded letter in her hands.
✓ The camera pans right with large amplitude at fast speed, revealing the open doorway.
✓ The camera holds a static shot as the runner exits the frame.
零件四:对白、旁白和演唱
说话的人要给一个稳定 ID,比如 (S1)、(S2)。多人同时说用 (S1,S2)。同一个人跨镜头保持同一个 ID,不出声的人不给 ID。
第一次出场时,要从画面和声音两个维度建立身份:年龄、性别、音高、音色、语速、口音。对白内容放在 标签里,格式是 [语言] 原文。原文一字不改,不翻译。
The young woman with a quiet, breathy voice (S1) says: [English] I get off at the next station.
The two children (S1,S2) shout together, [English] Wait for us!
旁白要写 says in an off-screen voiceover,而且每个旁白后面必须紧跟一句「画面里的人嘴唇没动」:
The man (S1) says in an off-screen voiceover: [English] I still remember that road. while his lips remain completely closed.
如果一句对白跨越了镜头切换,用 标记连接点,并写明声音跨镜头延续。如果对白被视频结尾截断,用 。
零件五:画面内文字
招牌、标签、字幕、霓虹灯——画面里观众能看到的文字,用英文双引号包裹,保留原文不翻译:
A red neon sign reading "营业中" glows above the doorway.
04 · 现场音和配乐:后两个部门
overall_soundscape:录音师的笔记
用 1-4 句英文写完整段视频的环境音、动作音和非语言人声。雨声、风声、车流、脚步、布料摩擦、碰撞、呼吸、笑声、喘气……全在这里。对白、唱歌和角色能听到的音乐不写这里。
overall_soundscape: Steady rain taps against the cafe windows while low room ambience continues underneath. The entrance bell rings once, followed by wet footsteps and the soft scrape of a chair.
如果整段视频完全静音,且用户明确要求了,才写 N/A。
non_diegetic_music:配乐监督的简报
1-3 句话描述只有观众能听到、角色听不到的背景音乐。聚焦乐器、速度、节奏、动态变化。别用「悲伤」「激昂」这种情绪词——模型需要的是具体的音乐描述。
non_diegetic_music: Sparse piano notes at a slow tempo, joined by sustained low strings that gradually increase in volume before fading out.
角色能听到的音乐(收音机、电视、手机播放的歌)不算配乐,那是剧情内声音,要写在 integrated_multimodal_description 里。没有背景音乐就写 N/A。
05 · 四个完整案例,逐行拆解
光说不练假把式。下面四个案例分别对应四种任务类型,每个都给完整提示词 + 逐行解读。
案例 1 · T2VA:清晨面包房
没有参考图,纯文字构建。你可以自由添加场景、人物和声音细节,只要跟用户意图一致。
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium-wide shot frames a baker opening the shutters of a small street bakery before sunrise. The camera pushes in with small amplitude at slow speed as the middle-aged baker with a calm, slightly raspy voice (S1) places a fresh loaf on the wooden counter and says: [English] First batch of the morning. [Shot 2] At 00:05.000, the camera cuts to a close-up of steam rising from the sliced bread while the baker's final words carry over from the previous shot.
overall_soundscape: Wooden shutters scrape open over a quiet street as trays clink softly inside the bakery. The doorbell rings once, followed by light footsteps and the crisp sound of bread being sliced.
non_diegetic_music: A soft acoustic-guitar pattern at a moderate tempo, joined by sparse upright-bass notes and a gentle fade at the end.
逐行拆解:
→ Shot 1 开头定了风格(真人实拍 + 电影感)和景别(中远景),接着交代主体(面包师)、动作(拉开卷帘门)、时间(日出前)。
→ 镜头运动:「缓慢小幅推进」——注意它写在动作中间,不是堆在句尾。
→ 说话人第一次出场就给了 ID (S1),并描述了音色(平静、略沙哑)。对白用 标签,标明语言 [English]。
→ Shot 2 在 5 秒处硬切,换成了面包切面的特写。最后一句「baker's final words carry over」说明对白声音跨镜头延续了。
→ soundscape 写了卷帘门声、托盘碰撞声、门铃声、脚步声、切面包声——全是环境音和动作音,没有对白。
→ music 写了具体乐器(原声吉他 + 低音提琴)、速度(中速)、结尾处理(渐弱)。
案例 2 · I2VA:雨夜列车
有一张首帧参考图。先写对齐指令,然后以图中的人物、构图、场景为起点,接着往下发展。
For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, the young woman shown in remains beside the rain-covered train window, preserving her appearance, clothing, seat position, and the carriage layout. The camera trucks right with small amplitude at slow speed as she lifts her gaze from the folded letter toward the passing city lights. Her reflection moves across the glass while the quiet, breathy young woman (S1) says: [English] I get off at the next station. She folds the letter along its existing crease.
overall_soundscape: The train wheels produce a steady metallic rhythm beneath a low ventilation hum. Rain ticks against the window while paper rustles softly in her hands.
non_diegetic_music: Sustained cello notes at a slow tempo with widely spaced piano tones, gradually decreasing in volume.
逐行拆解:
→ 第一行对齐指令:「目标视频 0.00 秒处,Picture 1(来自 Shot 1)被完全参考。」这是 I2VA 的固定句式。
→ Shot 1 先声明「图中年轻女子保持在雨中车窗旁」,并强调外貌、服装、座位、车厢布局全部保持一致——这就是 I2VA 的核心约束。
→ 推荐结构:首帧锚定 → 动作起始 → 连续发展 → 结果或反应。这里:锚定位置 → 抬头看窗外灯光 → 反射在玻璃上移动 → 说话 → 折信。
→ 镜头运动用了 trucks right with small amplitude at slow speed——机身向右缓慢小幅平移。
→ 声音层写了车轮金属节奏、通风嗡鸣、雨打窗户、纸张摩擦——全是列车场景该有的声音。
案例 3 · FL2VA:雨中开伞
有首尾两张图。重点不是重复描述两张静态图,而是写出从首帧到尾帧的运动路径。FL2VA 通常用单个镜头,让模型在首尾之间连续插值。这是一个 8 秒单镜头的例子。
How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot 1) aligns with the 8.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a rain-soaked cyclist begins in the position and framing established by Picture 1, holding a closed black umbrella beside a silver bicycle. The camera pulls out with small amplitude at slow speed as she releases the bicycle handle, raises the umbrella above her shoulder, and presses the runner upward until the canopy opens. Water rolls from the expanding fabric while she steps beneath it, rotates the handle into the final angle, and settles into the pose, spacing, and composition established by Picture 2 at the end of the shot.
overall_soundscape: Rain falls steadily on the pavement, followed by the metallic click of the umbrella runner and the soft snap of the canopy opening. Water drips from the bicycle frame as distant traffic passes.
non_diegetic_music: N/A
逐行拆解:
→ 对齐指令:Picture 1 对应 0 秒,Picture 2 对应 8 秒,都在 Shot 1 里——因为是单镜头。
→ 推荐结构:首帧状态 → 可观测的中间变化 → 差异逐步缩小 → 尾帧状态。这里:骑车人持闭合黑伞 → 松车把 → 举伞过肩 → 推伞骨 → 伞面撑开 → 水珠滚落 → 走到伞下 → 转手柄到最终角度 → 落入 Picture 2 的姿态。
→ 注意最后一句明确写了「settles into the pose, spacing, and composition established by Picture 2」——主动告诉模型尾帧长什么样。
→ 没有背景音乐,写 N/A。声音层全是现场动作音:雨声、伞骨金属咔嗒声、伞面撑开的啪声、水滴声、远处车流声。
案例 4 · L2VA:碎杯落地
只有一张尾帧图。你需要反推一个合理的前情,让动作自然地落到这张图上。这是一个 6 秒单镜头的例子。
How the reference pictures align with the target video — (from [Shot 1]) aligns with the 6.00-second mark of the target video.
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a close shot begins with an intact drinking glass near the edge of a dark wooden table, while the same hand and sleeve visible in approach from the right. The camera pushes in with small amplitude at slow speed as the fingertips strike the rim. The glass tips, falls, and hits the floor with a sharp impact; cracks spread through it as fragments slide outward. Toward the end, the moving pieces lose momentum and settle into the exact broken arrangement, hand position, camera angle, lighting, and final composition established by .
overall_soundscape: Fingertips tap the glass before it scrapes across the tabletop, falls, and breaks with a sharp crash. Small fragments scatter and gradually stop sliding across the floor.
non_diegetic_music: A low electronic pulse at a slow tempo, ending immediately after the glass breaks.
逐行拆解:
→ 对齐指令:Picture 1 对应 6 秒(视频结尾),属于 Shot 1 但不是开头——它是最终着陆点。
→ 推荐结构:合理的前情状态 → 明确动作和过渡路径 → 最终镜头逐步收敛 → 尾帧着陆。这里:完整杯子在桌边 → 手从右侧靠近 → 指尖碰杯沿 → 杯子翻倒 → 落地碎裂 → 裂纹扩散 → 碎片滑出 → 最终静止,落入 Picture 1 的碎片排列、手部位置、角度、光线和构图。
→ 关键技巧:开头建立了一个跟尾帧兼容但不相同的前情(完整杯子 vs 碎杯子),然后让动作一步步把差距缩小。
→ 最后一句把碎片排列、手部位置、镜头角度、光线、构图全部对齐到 Picture 1——越具体,模型越能精准着陆。
06 · 小白避坑指南
看完案例,你可能手痒想试了。先记住这几个常见坑:
坑一:对白不标语言。 标签里必须写 [English]、[Chinese] 等语言标记,否则模型可能乱发音。
坑二:镜头运动堆在句尾当标签。别写「... Zoom In, Push In, Pan Right」这种叠加标签。把镜头运动写成一句自然的话,融在动作描述里。
坑三:音乐写情绪不写乐器。「sad music」「epic soundtrack」这种词模型不认。写「sparse piano at slow tempo」「low strings gradually increasing in volume」。
坑四:I2VA 不写一致性约束。 图生视频时,必须在 Shot 1 明确写出「人物外貌、服装、座位、场景布局保持与 Picture 1 一致」,否则人物会变样。
坑五:对白和声音重复写。 对白只写在 integrated_multimodal_description 里,overall_soundscape 只写环境音和动作音。别在 soundscape 里重复对白内容。
坑六:FL2VA 写成两张静态图的描述。 首尾帧任务的重点是中间的运动路径,不是重复描述首帧和尾帧各自长什么样。
07 · 系统提示词:让 AI 帮你写 H3 提示词
如果你觉得手写还是太麻烦,这里有一段系统提示词。把它发给任何大语言模型(GPT、Claude、DeepSeek、Kimi 都行),再告诉它你想要什么视频,它就会按 H3 的格式帮你生成标准提示词。
你是一个 MiniMax-H3 视频生成提示词专家。你的任务是根据用户的描述,生成符合 MiniMax-H3 格式要求的视频提示词。
## 任务类型
根据用户提供的素材判断任务类型:
- T2VA:纯文本生成视频,无参考图。提示词直接从三个核心字段开始。
- I2VA:一张首帧参考图。提示词第一行为对齐指令,然后空一行,再写三个核心字段。
- FL2VA:首帧+尾帧两张参考图。提示词第一行为对齐指令,然后空一行,再写三个核心字段。
- L2VA:一张尾帧参考图。提示词第一行为对齐指令,然后空一行,再写三个核心字段。
## 对齐指令模板
- I2VA: For the target video, at 0.00 seconds into the target video, (from [Shot 1]) is fully referenced.
- FL2VA(N 为实际最后一个镜头编号,S.SS 为视频时长,保留两位小数): How the reference pictures align with the target video — Picture 1 (from Shot 1) aligns with the 0.00-second mark of the target video; Picture 2 (from Shot N) aligns with the S.SS-second mark of the target video.
- L2VA(N 为实际最后一个镜头编号,S.SS 为视频时长,保留两位小数): How the reference pictures align with the target video — (from [Shot N]) aligns with the S.SS-second mark of the target video.
## 三个核心字段
1. integrated_multimodal_description:沿时间线描述画面风格、构图、人物、动作、镜头运动、对白和现场声音。
2. overall_soundscape:1-4 句英文,概括全视频的环境音、动作音和非语言人声。对白和唱歌不写这里。
3. non_diegetic_music:1-3 句英文,描述只有观众能听到的背景音乐,聚焦乐器、速度、节奏和动态变化。无配乐写 N/A。
## 写作规则
1. [Shot 1] 开头先定风格(Cinematic/live-action/2D-animated/3D CG/claymation/watercolor/vintage film)和初始构图。
2. 第一个镜头不加时间戳;后续镜头用 [Shot N] At MM:SS.mmm, the camera cuts to... 格式,时间严格递增。
3. 镜头运动写成自然英文动作句,格式为:运动类型 + with small/large amplitude + at slow/fast speed。中等幅度和正常速度可省略。
4. 说话人用稳定 ID (S1)(S2),多人同说用 (S1,S2)。首次出场需描述年龄、性别、音色等身份信息。对白格式:[语言] 原文内容,原文一字不改不翻译。
5. 旁白用 says in an off-screen voiceover,其后必须写明画面中人物嘴唇保持闭合。
6. 对白跨镜头用 标记并写明声音延续;被结尾截断用 。
7. 画面内可见文字(招牌、字幕、标签)用英文双引号包裹,保留原文不翻译。
8. I2VA:Shot 1 须声明人物外貌、服装、位置、场景与 Picture 1 保持一致,然后从首帧发展动作。
9. FL2VA:通常用单镜头,写从首帧到尾帧的运动路径。末尾须明确对齐到 Picture 2。
10. L2VA:反推兼容的前情状态,让动作逐步收敛到尾帧。末尾须明确对齐到 Picture 1。
11. 整个提示词用英文撰写( 内的对白原文和画面内文字保留原始语言)。
## 输出格式
直接输出完整提示词,不要加任何解释、注释或 markdown 代码块标记。
## 用户输入
{用户会在这里描述想要的视频内容、提供参考图信息、指定视频时长等}
使用方法很简单:
第一步:把上面框里的系统提示词复制发给 AI。
第二步:告诉 AI 你想要什么。比如:「帮我写一个 I2VA 提示词,参考图是一个穿红色卫衣的年轻男子站在天台上,夕阳逆光,视频时长 6 秒,他回头看了一眼城市天际线,说了句'该走了'。」
第三步:AI 会输出标准格式的提示词,你复制粘贴到 H3 的输入框,配上参考图,点生成。
写好提示词,你就已经赢了一半
MiniMax-H3 是一个听指令干活的全能剧组。你给它的指令越像一份专业的镜头脚本,它给你的回报就越像一段专业的短片。
记住三个核心:选对任务类型,写满三个字段,把镜头运动和对白写具体。
剩下的,交给剧组。