更新于 2026-09-30:本页已按本次最终 v4 成片重整。当前基线是 13 段国外电影、98.5 秒、1280×720、30fps;画面等比放大并裁切铺满,左上角片名使用纯黑矩形底与白字,中英字幕只对应演员真实对白。原来的 61 秒 / 10 段方案是历史版本,下面的代码与参数以最终版本为准。
这份手册面向熟悉基本终端操作、希望复用 Python 与 FFmpeg 做电影混剪的人。重点不仅是把视频拼起来,还要让镜头、对白字幕、人声和音乐始终对应:没有明确台词的镜头可以留白,有对白字幕的段落就应当听得到对应的人声,片尾音乐必须陪画面一起淡出。
一、最终成片规则:先把审美要求写成可执行配置
| 项目 | 最终约定 |
|---|---|
| 片段 | 13 段;每段 8–9 秒,仍在 5–10 秒范围内 |
| 画幅 | 1280×720;等比放大覆盖画布后居中裁切,不主动补上下黑边 |
| 片名 | 左上角“电影名 年份”;纯黑不透明矩形背景、白字、无竖线、无描边 |
| 字幕 | 中文主字幕与英文副字幕,只保留演员真实对白;3 段没有对白字幕 |
| 源音频 | 默认不混入原音;10 个有明确台词的片段仅使用 Demucs 分离后、按句截取的人声 |
| BGM | Kevin MacLeod《At Rest》;源曲从 44 秒取用,成片 3.5 秒起渐入,一首曲子连续贯穿 |
| 转场 | 30fps 下 20 帧交叉溶解,即 0.6667 秒 |
| 结尾 | 最后 3 秒画面、文字与音乐同步淡出至黑 |
| 交付 | MP4、SRT、来源清单、独立人声、JSON 时间轴、分阶段脚本、验收图与桌面软链接 |
“默认静音”与“有台词的片段保留人声”并不矛盾:默认移除原片混合音轨;决定展示某句真实对白时,再单独取它的人声。不要把它解释成“只放字幕也算完成对白”。这次初版只留了四处人声,六个明确台词片段只有字幕,用户反馈后才补齐。
没有对白也不需要硬填文字。《迷失东京》《爱乐之城》《E.T.》最终只保留画面、片名和音乐;初版添加的原创旁白已全部删掉。拥抱、回望、微笑本身就能表达离别,留白也是节奏的一部分。
二、片单、节奏与时间轴
本次的情绪线是恋人告别 → 亲情与生死 → 感谢、成长与祝福。影片年份写在画面里,片段顺序按情绪组织,不要求按年代排序。
| 序号 | 电影 | 成片时间 / 秒 | 源入点 / 秒 | 声音 |
|---|---|---|---|---|
| 1 | Casablanca 1942 | 0.00–8.50 | 44.80 | 分离人声 |
| 2 | Roman Holiday 1953 | 7.83–15.83 | 154.00 | 分离人声 |
| 3 | Before Sunrise 1995 | 15.17–23.17 | 124.80 | 分离人声 |
| 4 | Lost in Translation 2003 | 22.50–30.50 | 160.00 | 仅 BGM |
| 5 | Brokeback Mountain 2005 | 29.83–37.83 | 236.70 | 分离人声 |
| 6 | La La Land 2016 | 37.17–45.17 | 543.00 | 仅 BGM |
| 7 | Interstellar 2014 | 44.50–53.50 | 189.25 | 分离人声 |
| 8 | Titanic 1997 | 52.83–60.83 | 281.80 | 分离人声 |
| 9 | Forrest Gump 1994 | 60.17–68.17 | 168.50 | 分离人声 |
| 10 | Cinema Paradiso 1988 | 67.50–75.50 | 26.00 | 分离人声 |
| 11 | Good Will Hunting 1997 | 74.83–82.83 | 1.00 | 分离人声 |
| 12 | E.T. 1982 | 82.17–90.17 | 88.00 | 仅 BGM |
| 13 | Toy Story 3 2010 | 89.50–98.50 | 40.50 | 分离人声 |
《星际穿越》由 8 秒延长到 9 秒,让“I’m coming back”完整说完。因此最终总长是 98.5 秒,不能继续沿用初版 97.5 秒,也不能使用旧手册的 61 秒固定公式。
FPS = 30
requested_transition = 0.65
transition_frames = round(requested_transition * FPS) # 20
transition = transition_frames / FPS # 0.666666...
durations = [8.5, 8, 8, 8, 8, 8, 9, 8, 8, 8, 8, 8, 9]
starts, cursor = [], 0.0
for i, duration in enumerate(durations):
starts.append(cursor)
cursor += duration - (transition if i < len(durations) - 1 else 0)
print(cursor) # 98.5
所有时间先落到整数帧,片段起点逐段累加。0.65 秒在 30fps 下是 19.5 帧,不能把它原样当成最终精确时长。字幕与人声也使用同一份配置,避免各自维护一套时间码。
三、环境与项目目录
brew install ffmpeg python@3.13 yt-dlp
mkdir -p ~/Desktop/farewell-mashup/{clips,bgm,output,work,docs}
cd ~/Desktop/farewell-mashup
python3.13 -m venv work/venv
./work/venv/bin/pip install pillow numpy faster-whisper demucs
./work/venv/bin/pip freeze > work/requirements-lock.txt
ffmpeg -hide_banner -filters | rg 'xfade|overlay|loudnorm|sidechaincompress|alimiter'
将本文后面的脚本保存到 work/ 下,原始片段放 clips/,音乐放 bgm/。Python 虚拟环境和模型首次准备需要下载;换机器时重建虚拟环境,不要直接复制旧 venv 后假设它仍可运行。版本以本项目的 requirements-lock.txt 为准。
av==19.0.0
ctranslate2==4.8.2
demucs==4.1.0
faster-whisper==1.2.1
numpy==2.5.3
pillow==12.3.0
torch==2.14.0
farewell-mashup/
├── clips/ # 完整下载片段,便于重新选点
├── bgm/ # At Rest - Kevin MacLeod.mp3 与署名
├── output/ # 最新成片、SRT、验收报告与核验图
├── docs/ # 来源、修订记录、复用说明
└── work/
├── timeline.json # 唯一编排配置
├── build.py # segments / voices / join / titles / mix
├── probe.py # 区间对白定位
├── verify.py # 编码、时长、峰值、黑场、抽帧
├── check_audio_tail.py # 最后三秒音乐连续性回归检查
├── segments/ # 无声统一分段与抽取的原音
├── separated/ # Demucs 缓存
├── voices.wav # 对齐总时间轴的独立人声
├── picture.mp4 # 画面转场后、无文字
├── overlays.mov # 有限帧数透明文字轨
└── captioned.mp4 # 烧录字幕后的画面
本机 FFmpeg 没有 drawtext / ass / subtitles 滤镜,因此文字使用 Pillow 绘制,再叠到视频上。代码使用 macOS 自带 Songti 与 Arial;跨平台先配置可用的中文字体,不能用缺少中文字形的字体静默兜底。
四、选素材:检索、抽帧与转写三者一起核对
- 检索结果不等于正确素材。本次《天堂电影院》的首个结果是片尾字幕,已经换成实际告别场景。记录来源链接、电影年份、文件名与实际入点。
- 先抽帧查看场景,再在候选对白附近做短区间转写。字幕不可凭记忆填写,自动转写也可能把 Paris、Jenny 等词识别错。
- 稀疏、轻声对白可能被 VAD 漏掉。《玩具总动员3》在全段 VAD 转写中漏掉“So long, partner”;关闭 VAD 后才定位到源片约 42.82–44.58 秒。
- 对听不清的耳语留空;不要把自己的解读写成演员说过的词。最终版本也不再用原创字幕补满无对白镜头。
- cropdetect 可以辅助查源黑边,但暗场景会被误判为黑边。《爱乐之城》的暗背景曾导致过度裁切建议;必须用抽帧核对再填 crop。
音频输入统一由 FFmpeg 解码为 16kHz float32 PCM,再以 numpy 数组喂给 faster-whisper。本次 PyAV 与 faster-whisper 的 metadata_errors 参数不兼容,直接传文件路径报错;PCM 路径避开了这个编解码库接口问题。
完整 probe.py:定位真实对白
import sys,subprocess,numpy as np
from faster_whisper import WhisperModel
p,a,b=sys.argv[1],float(sys.argv[2]),float(sys.argv[3]); raw=subprocess.check_output(['ffmpeg','-v','error','-ss',str(a),'-t',str(b-a),'-i',p,'-f','f32le','-ar','16000','-ac','1','-'])
m=WhisperModel('small',device='cpu',compute_type='int8')
s,info=m.transcribe(np.frombuffer(raw,dtype=np.float32),language='en',vad_filter=False,word_timestamps=True)
for x in s: print(a+x.start,a+x.end,x.text,flush=True)
cd ~/Desktop/farewell-mashup/work
./venv/bin/python probe.py ../clips/07-titanic.mp4 273 296
# 实测目标:284.00–284.70 I'll never let go.
# 285.38–286.14 I promise.
五、画幅与文字:去掉重复黑边,字幕不参与转场
本次最初为文字预留了上下黑区,用户希望画面铺满。最终改为先按配置去除源黑边或硬字幕区,再等比放大覆盖画布,最后居中裁切。人物比例不变,但会裁掉一部分边缘;要逐镜头检查脸、手与重要道具,必要时针对该镜头调整构图。不要把低清素材放大后称作原生高清。
scale=1280:720:force_original_aspect_ratio=increase:flags=lanczos,
crop=1280:720:(iw-ow)/2:(ih-oh)/2,
setsar=1,fps=30,settb=AVTB,setpts=PTS-STARTPTS,format=yuv420p
片名采用不透明纯黑矩形背景与白字,黑底覆盖整个标签区域,并留内边距。不要用黑色文字描边替代黑底,也不加装饰竖线。演员对白则显示在画面底部安全区,用描边适应明亮背景。
title = f"{s['film']} {s['year']}"
ft = font(26, True)
box = d.textbbox((0, 0), title, font=ft)
tw, th = box[2] - box[0], box[3] - box[1]
d.rectangle((24, 20, 24 + tw + 24, 20 + th + 20), fill=(0, 0, 0, 255))
d.text((36 - box[0], 30 - box[1]), title, font=ft, fill=(255, 255, 255, 255))
先拼干净画面,再统一叠文字。相邻段字幕窗口避开溶解区,片名和对白分别淡入淡出,不能让片名跟随每一句字幕一起闪烁。最终采用有限帧数的 QTRLE 透明 MOV,而非几十路无限 loop 的 PNG:输入、输出帧数明确,更容易防止无限空转和字幕只闪一帧。
六、音频:字幕、人声与音乐必须一起设计
每句对白建立 voices 时间窗,Demucs 分离 vocals 后仅剪入该句;所有未选中的原音不进入混音。对白边缘加约 40ms 淡入淡出,避免切口噪声。人声先归一化、再对齐主时间轴;BGM 只处理一次,不能每个镜头重新起音乐。
Demucs 能减少原片配乐,但不保证完全移除混响、环境声或所有残留。必须检查分离结果;不能把 highpass / lowpass 均衡滤波当成人声分离。清晰度不够时应换素材、换镜头或调整分离与混音,而不是把原片整条混合音轨加回来。
本次 BGM 选择《At Rest》,从源曲 44 秒取用,成片 3.5 秒起渐入。对白触发侧链压缩,音乐随说话适度降低、结束后恢复;音乐不能在对白处硬切静音。选曲也不能只找 RMS 最大的高潮,题材、旋律、情绪推进和尾部余韵同样影响是否合适。
6.1 片尾音乐突然消失:不是文件时长问题,而是时间戳问题
初版视频有 97.5 秒,但真正的有声样本只到约 94.02 秒,后面约 3.5 秒全是静音。原曲与单独截取的 94 秒音乐都完整,因此问题不在素材。加上 3.5 秒 adelay 后没有重置时间戳,后续 atrim 提前裁掉了尾音;apad 又用静音补足文件,导致只看总时长和峰值的验收误判为通过。
| 检查对象 | 实测结果 |
|---|---|
| 原曲 | 解码 206.568 秒,尾部仍有声 |
| 仅截取 start=44、duration=94 | 94 秒完整有声 |
| 初版最终音频 | 文件约 97.5 秒,有声只到约 94.02 秒 |
| 修复后延迟音乐 | 94 秒音乐 + 3.5 秒入场延迟 = 97.5 秒,尾音保留 |
# 错误链路:延迟之后立刻按偏移后的时间戳裁切
adelay=3500|3500,apad,atrim=duration=97.5
# 本次实测修复:延迟之后,按样本数重建从0开始的时间戳
adelay=3500|3500,asetpts=N/SR/TB,apad,atrim=duration=97.5
修复后按当前 98.5 秒时间轴重新混音,最后三秒逐渐淡出。除了总文件时长,还要检查淡出前半段、后半段、最后 0.4 秒的短窗口,不能出现提前归零的静音段。最后一帧接近黑、最后音量趋近零是正确淡出;提前数秒完全静音不是。
6.2 本次已通过的尾音回归检查
| 时间窗口 / 秒 | 修订版单声道 RMS |
|---|---|
| 95.5–96.0 | 0.0264518 |
| 96.5–97.0 | 0.0120220 |
| 97.5–98.0 | 0.0112467 |
| 98.1–98.4 | 0.0019714 |
这些数字用于证明尾部持续有声,不代表每个短窗口必须严格单调下降;乐曲本身的音符强弱会变化。测试阈值针对本次曲目,换成有自然静音的音乐时应重新判断,不能机械套用。
完整 check_audio_tail.py:防止静音填充掩盖尾音裁切
import pathlib,subprocess,numpy as np,sys,json
r=pathlib.Path(__file__).resolve().parents[1];p=pathlib.Path(sys.argv[1]) if len(sys.argv)>1 else r/'output/离别_Farewell.mp4'
a=np.frombuffer(subprocess.check_output(['ffmpeg','-v','error','-i',str(p),'-vn','-ar','48000','-ac','1','-f','f32le','-']),np.float32)
total=float(subprocess.check_output(['ffprobe','-v','error','-show_entries','format=duration','-of','csv=p=0',str(p)]))
for lo,hi in [(total-3,total-2.5),(total-2,total-1.5),(total-1,total-.5),(total-.4,total-.1)]:
b=a[round(lo*48000):round(hi*48000)];rms=float(np.sqrt(np.mean(b*b)));print(lo,hi,rms);assert rms>1e-5,f'BGM lost before fade ends at {lo}s'
print('PASS: music remains audible until the end of the fade')
七、完整配置与主脚本:换主题优先改配置
把下面 JSON 保存为 work/timeline.json。inpoint 是源视频入点,cues 是该段内部的字幕时间,voices 是该段内部的人声窗口;空 cues 与空 voices 表示只留画面和音乐。crop 必须针对实际下载素材核对,不同上传版本不能直接套用同一个入点。
完整 timeline.json:13 段最终编排
{
"width": 1280,
"height": 720,
"fps": 30,
"transition": 0.65,
"bgm": "At Rest - Kevin MacLeod.mp3",
"bgm_start": 44,
"bgm_delay": 3.5,
"end_fade": 3,
"segments": [
{
"source": "01-casablanca.mp4",
"film": "Casablanca",
"year": 1942,
"inpoint": 44.8,
"duration": 8.5,
"cues": [
{
"start": 0.55,
"end": 2.15,
"zh": "那我们呢?",
"en": "But what about us?",
"kind": "dialogue"
},
{
"start": 2.65,
"end": 4.35,
"zh": "我们永远拥有巴黎。",
"en": "We'll always have Paris.",
"kind": "dialogue"
}
],
"voices": [
[
0.45,
4.25
]
],
"crop": "652:480:98:0"
},
{
"source": "02-romanholiday.mp4",
"film": "Roman Holiday",
"year": 1953,
"inpoint": 154,
"duration": 8,
"cues": [
{
"start": 0.8,
"end": 7.2,
"zh": "我会珍藏在这里的回忆,直至生命尽头。",
"en": "I will cherish my visit here in memory as long as I live.",
"kind": "dialogue"
}
],
"voices": [
[
0.65,
7.2
]
],
"crop": null
},
{
"source": "06-beforesunrise.mp4",
"film": "Before Sunrise",
"year": 1995,
"inpoint": 124.8,
"duration": 8,
"cues": [
{
"start": 1.5,
"end": 5.8,
"zh": "再见。",
"en": "Goodbye.",
"kind": "dialogue"
}
],
"voices": [
[
1.7,
7.05
]
],
"crop": "iw:ih*0.84:0:0"
},
{
"source": "08-lostintranslation.mp4",
"film": "Lost in Translation",
"year": 2003,
"inpoint": 160,
"duration": 8,
"cues": [],
"voices": [],
"crop": null
},
{
"source": "09-brokeback.mp4",
"film": "Brokeback Mountain",
"year": 2005,
"inpoint": 236.7,
"duration": 8,
"cues": [
{
"start": 1.1,
"end": 4.9,
"zh": "我真希望知道,该如何戒掉你。",
"en": "I wish I knew how to quit you.",
"kind": "dialogue"
}
],
"voices": [
[
1.2,
4.8
]
],
"crop": "1920:1040:0:20"
},
{
"source": "11-lalaland.mp4",
"film": "La La Land",
"year": 2016,
"inpoint": 543,
"duration": 8,
"cues": [],
"voices": [],
"crop": "1920:752:0:164"
},
{
"source": "10-interstellar.mp4",
"film": "Interstellar",
"year": 2014,
"inpoint": 189.25,
"duration": 9,
"cues": [
{
"start": 1.7,
"end": 3,
"zh": "我爱你。",
"en": "I love you.",
"kind": "dialogue"
},
{
"start": 6.4,
"end": 8.2,
"zh": "我会回来的。",
"en": "And I'm coming back.",
"kind": "dialogue"
}
],
"voices": [
[
1.65,
2.75
],
[
6.4,
8.25
]
],
"crop": null
},
{
"source": "07-titanic.mp4",
"film": "Titanic",
"year": 1997,
"inpoint": 281.8,
"duration": 8,
"cues": [
{
"start": 2.1,
"end": 3.45,
"zh": "我永不放手。",
"en": "I'll never let go.",
"kind": "dialogue"
},
{
"start": 3.5,
"end": 5.4,
"zh": "我保证。",
"en": "I promise.",
"kind": "dialogue"
}
],
"voices": [
[
2,
4.5
]
],
"crop": "1920:816:0:132"
},
{
"source": "05-forrestgump.mp4",
"film": "Forrest Gump",
"year": 1994,
"inpoint": 168.5,
"duration": 8,
"cues": [
{
"start": 1.1,
"end": 4.5,
"zh": "我想你,珍妮。",
"en": "I miss you, Jenny.",
"kind": "dialogue"
}
],
"voices": [
[
1.25,
3.25
]
],
"crop": null
},
{
"source": "04b-cinemaparadiso.mp4",
"film": "Cinema Paradiso",
"year": 1988,
"inpoint": 26,
"duration": 8,
"cues": [
{
"start": 3.2,
"end": 6.9,
"zh": "无论做什么,都要热爱它。",
"en": "Whatever you end up doing, love it.",
"kind": "dialogue"
}
],
"voices": [
[
3,
7.2
]
],
"crop": "iw:ih*0.70:0:0"
},
{
"source": "12-goodwillhunting.mp4",
"film": "Good Will Hunting",
"year": 1997,
"inpoint": 1,
"duration": 8,
"cues": [
{
"start": 3.1,
"end": 6.25,
"zh": "抱歉,我得去找一个姑娘。",
"en": "Sorry, I had to go see about a girl.",
"kind": "dialogue"
}
],
"voices": [
[
3,
6.15
]
],
"crop": "iw*0.76:ih*0.65:iw*0.12:ih*0.12"
},
{
"source": "03-et.mp4",
"film": "E.T.",
"year": 1982,
"inpoint": 88,
"duration": 8,
"cues": [],
"voices": [],
"crop": "1280:688:0:16"
},
{
"source": "13-toystory.mp4",
"film": "Toy Story 3",
"year": 2010,
"inpoint": 40.5,
"duration": 9,
"cues": [
{
"start": 2.1,
"end": 4.7,
"zh": "再见,老伙计。",
"en": "So long, partner.",
"kind": "dialogue"
}
],
"voices": [
[
2.2,
4.4
]
],
"crop": null
}
],
"framing": "cover"
}
将下面主脚本保存为 work/build.py。它以脚本所在位置解析项目根目录,不硬编码用户名与桌面目录。构建阶段分开,改字幕不用重跑人声分离,改音乐不用重编码视频。
完整 build.py:分段、人声、转场、文字与混音
#!/usr/bin/env python3
"""Portable, config-driven movie montage. Run: python build.py [segments|voices|join|titles|mix|all]."""
import pathlib,json,subprocess,sys,math
import numpy as np
from PIL import Image,ImageDraw,ImageFont
R=pathlib.Path(__file__).resolve().parents[1]; WK=R/'work'; C=json.loads((WK/'timeline.json').read_text()); S=C['segments']; W,H,FPS=C['width'],C['height'],C['fps']; XF=round(C['transition']*FPS)/FPS
for d in ['segments','overlays','logs']: (WK/d).mkdir(exist_ok=True)
clock=0
for i,s in enumerate(S):
s['duration']=round(s['duration']*FPS)/FPS;s['start']=clock;clock+=s['duration']-(XF if i<len(S)-1 else 0)
TOTAL=round(clock*FPS)/FPS;FRAMES=round(TOTAL*FPS)
def run(args,name):
print(name,flush=True)
with open(WK/'logs'/(name+'.log'),'w') as log: subprocess.run(['ffmpeg','-hide_banner','-v','warning','-y',*map(str,args)],stdout=log,stderr=log,check=True)
def enc():return ['-c:v','libx264','-preset','fast','-crf','18','-pix_fmt','yuv420p','-threads','2']
def segments():
for i,s in enumerate(S):
dst=WK/'segments'/f'{i:02}.mp4'
filt=[]
if s.get('crop'):filt+=['crop='+s['crop']]
filt+=[f'scale={W}:{H}:force_original_aspect_ratio=increase:flags=lanczos',f'crop={W}:{H}:(iw-ow)/2:(ih-oh)/2','setsar=1',f'fps={FPS}','settb=AVTB','setpts=PTS-STARTPTS','format=yuv420p']
run(['-ss',s['inpoint'],'-i',R/'clips'/s['source'],'-t',s['duration'],'-an','-vf',','.join(filt),*enc(),'-frames:v',round(s['duration']*FPS),dst],f'segment-{i:02}')
def voices():
sr=48000; master=np.zeros((round(TOTAL*sr),2),np.float32)
for i,s in enumerate(S):
if not s['voices']:continue
raw=WK/'segments'/f'voice-raw-{i:02}.wav';stem=WK/'separated'/'htdemucs'/raw.stem/'vocals.wav'
if not stem.exists():
run(['-ss',s['inpoint'],'-i',R/'clips'/s['source'],'-t',s['duration'],'-vn','-ar','44100','-ac','2',raw],f'voice-extract-{i:02}')
print('demucs',i,flush=True)
with open(WK/'logs'/f'demucs-{i:02}.log','w') as log:subprocess.run([sys.executable,'-m','demucs','--two-stems','vocals','-n','htdemucs','--shifts','0','-d','cpu','-o',str(WK/'separated'),str(raw)],stdout=log,stderr=log,check=True)
pcm=subprocess.check_output(['ffmpeg','-v','error','-i',str(stem),'-af','highpass=f=100,lowpass=f=8000,loudnorm=I=-18:TP=-3:LRA=7','-f','f32le','-ar',str(sr),'-ac','2','-'])
a=np.frombuffer(pcm,np.float32).reshape(-1,2)
for v0,v1 in s['voices']:
l,h=round(v0*sr),min(round(v1*sr),len(a));b=a[l:h].copy();n=min(round(.04*sr),len(b)//2);b[:n]*=np.linspace(0,1,n)[:,None];b[-n:]*=np.linspace(1,0,n)[:,None]
pos=round((s['start']+v0)*sr);master[pos:pos+len(b)]+=b
print('voice isolated',s['film'],flush=True)
p=subprocess.run(['ffmpeg','-v','error','-y','-f','f32le','-ar',str(sr),'-ac','2','-i','-','-c:a','pcm_s24le',str(WK/'voices.wav')],input=master.tobytes(),check=True)
def join():
args=[]
for i in range(len(S)):args+=['-i',WK/'segments'/f'{i:02}.mp4']
filters=[];prev='0:v'
for i in range(1,len(S)):
lab=f'x{i}';filters.append(f'[{prev}][{i}:v]xfade=transition=fade:duration={XF}:offset={S[i]["start"]:.6f}[{lab}]');prev=lab
run([*args,'-filter_complex_threads','1','-filter_complex',';'.join(filters),'-map',f'[{prev}]',*enc(),'-frames:v',FRAMES,WK/'picture.mp4'],'join')
def font(size,english=False):return ImageFont.truetype('/System/Library/Fonts/Supplemental/Arial.ttf' if english else '/System/Library/Fonts/Supplemental/Songti.ttc',size)
def textcenter(d,y,text,f,fill):
box=d.textbbox((0,0),text,font=f);x=(W-(box[2]-box[0]))/2;d.text((x,y),text,font=f,fill=fill,stroke_width=2,stroke_fill=(0,0,0,230))
def overlay_img(s,cue=None):
im=Image.new('RGBA',(W,H));d=ImageDraw.Draw(im)
if cue is None:
title=f'{s["film"]} {s["year"]}';ft=font(26,True);box=d.textbbox((0,0),title,font=ft)
tw,th=box[2]-box[0],box[3]-box[1]
d.rectangle((24,20,24+tw+24,20+th+20),fill=(0,0,0,255))
d.text((36-box[0],30-box[1]),title,font=ft,fill=(255,255,255,255))
if cue:
textcenter(d,610,cue['zh'],font(31),(250,250,250,255));textcenter(d,656,cue['en'],font(21,True),(206,206,206,255))
return im
def titles():
# One finite alpha movie prevents dozens of looped PNG inputs and never overlaps title cards.
events=[];srt=[];k=1
for i,s in enumerate(S):
lo=s['start']+(XF+.10 if i else 0);hi=s['start']+s['duration']-(XF+.10 if i<len(S)-1 else .15)
events.append((lo,hi,overlay_img(s)))
for q in s['cues']:
a=max(lo,s['start']+q['start']);b=min(hi,s['start']+q['end']);events.append((a,b,overlay_img(s,q)))
def tc(x):
n=round(x*1000);return f'{n//3600000:02}:{n//60000%60:02}:{n//1000%60:02},{n%1000:03}'
prefix=''
srt.append(f'{k}\n{tc(a)} --> {tc(b)}\n{prefix}{q["zh"]}\n{q["en"]}\n');k+=1
(R/'output'/'离别_中英字幕.srt').write_text('\n'.join(srt),encoding='utf-8')
blank=Image.new('RGBA',(W,H));blankbytes=blank.tobytes();cache={};last=None
with open(WK/'logs'/'overlay.log','w') as log:
p=subprocess.Popen(['ffmpeg','-v','warning','-y','-f','rawvideo','-pixel_format','rgba','-video_size',f'{W}x{H}','-framerate',str(FPS),'-i','-','-an','-c:v','qtrle','-pix_fmt','argb','-frames:v',str(FRAMES),str(WK/'overlays.mov')],stdin=subprocess.PIPE,stderr=log)
try:
for f in range(FRAMES):
t=f/FPS;active=[e for e in events if e[0]<=t<e[1]]
if active:
opacities=[max(0,min(1,(t-a)/.12,(b-t)/.12)) for a,b,im in active]
key=tuple((id(e[2]),round(alpha*12)) for e,alpha in zip(active,opacities))
if key not in cache:
frame=blank.copy()
for (a,b,im),opacity in zip(active,opacities):
layer=im.copy();layer.putalpha(im.getchannel('A').point(lambda z:int(z*opacity)))
frame=Image.alpha_composite(frame,layer)
cache[key]=frame.tobytes()
data=cache[key]
else:data=blankbytes
p.stdin.write(data)
p.stdin.close();assert p.wait()==0
except: p.kill();raise
run(['-i',WK/'picture.mp4','-i',WK/'overlays.mov','-filter_complex_threads','1','-filter_complex',f'[0:v][1:v]overlay=shortest=1,fade=t=out:st={TOTAL-C["end_fade"]}:d={C["end_fade"]}[v]','-map','[v]',*enc(),'-frames:v',FRAMES,WK/'captioned.mp4'],'titles')
def mix():
delay=C['bgm_delay'];fade=C['end_fade']
filt=f'[0:a]atrim=duration={TOTAL},asetpts=PTS-STARTPTS,asplit=2[voice][side];[1:a]atrim=start={C["bgm_start"]}:duration={TOTAL-delay},asetpts=PTS-STARTPTS,loudnorm=I=-21:TP=-4:LRA=9,afade=t=in:d=2,adelay={round(delay*1000)}|{round(delay*1000)},asetpts=N/SR/TB,apad,atrim=duration={TOTAL}[music];[music][side]sidechaincompress=threshold=0.025:ratio=5:attack=35:release=450[duck];[duck][voice]amix=inputs=2:normalize=0,alimiter=limit=0.92:level=false,afade=t=out:st={TOTAL-fade}:d={fade}[mix]'
run(['-i',WK/'voices.wav','-i',R/'bgm'/C['bgm'],'-i',WK/'captioned.mp4','-filter_complex',filt,'-map','2:v:0','-map','[mix]','-c:v','copy','-c:a','aac','-b:a','256k','-ar','48000','-t',TOTAL,'-movflags','+faststart',R/'output'/'离别_Farewell.mp4'],'mix')
(WK/'resolved_timeline.json').write_text(json.dumps(dict(total=TOTAL,transition=XF,segments=S),ensure_ascii=False,indent=2))
print('FINISHED',TOTAL,flush=True)
if __name__=='__main__':
stage=sys.argv[1] if len(sys.argv)>1 else 'all'
for name,func in [('segments',segments),('voices',voices),('join',join),('titles',titles),('mix',mix)]:
if stage in ['all',name]:func()
缓存约束:当前人声缓存按 voice-raw-NN 命名。修改某段源文件、入点或时长后,应清理该段 separated/htdemucs/voice-raw-NN,再重跑 voices;否则可能把旧声音配到新镜头。后续改进可用源文件指纹、入点、时长、模型与处理参数共同组成缓存键,这仍是待实现的增强,不是当前脚本已有能力。
八、运行与验收:验收对象必须是最终输出
cd ~/Desktop/farewell-mashup/work
./venv/bin/python build.py all
./venv/bin/python verify.py
./venv/bin/python check_audio_tail.py
# 仅改片名/字幕
./venv/bin/python build.py titles
./venv/bin/python build.py mix
# 仅改音乐/音量
./venv/bin/python build.py mix
# 仅改画幅,不改入点和时长
./venv/bin/python build.py segments
./venv/bin/python build.py join
./venv/bin/python build.py titles
./venv/bin/python build.py mix
完整 verify.py:参数、黑场、峰值与逐段抽帧
import json,pathlib,subprocess,numpy as np
from PIL import Image,ImageDraw,ImageFont
r=pathlib.Path(__file__).resolve().parents[1];c=json.loads((r/'work/timeline.json').read_text());d=json.loads((r/'work/resolved_timeline.json').read_text());out=r/'output/离别_Farewell.mp4';report={}
info=json.loads(subprocess.check_output(['ffprobe','-v','error','-show_streams','-show_format','-of','json',str(out)]));report['duration']=float(info['format']['duration']);report['size']=int(info['format']['size']);report['expected_duration']=d['total'];report['streams']=[{k:s.get(k) for k in ['codec_name','width','height','r_frame_rate','sample_rate','channels']} for s in info['streams']]
assert abs(report['duration']-d['total'])<.08
assert all(5<=s['duration']<=10 for s in d['segments'])
font=ImageFont.truetype('/System/Library/Fonts/Supplemental/Arial.ttf',16)
points=[(s['start']+min(3.3,s['duration']/2),s['film']) for s in d['segments']]
points += [(s['start']+d['transition']/2,'TRANSITION '+str(i)) for i,s in enumerate(d['segments'][1:],1)]
im=Image.new('RGB',(1280,mathrows:=((len(points)+3)//4)*205))
for i,(t,label) in enumerate(points):
p=r/'work'/f'qa-{i:02}.jpg';subprocess.run(['ffmpeg','-v','error','-y','-ss',str(t),'-i',str(out),'-frames:v','1','-vf','scale=320:180',str(p)],check=True)
im.paste(Image.open(p),(i%4*320,i//4*205));ImageDraw.Draw(im).text((i%4*320+4,i//4*205+182),f'{t:.2f} {label}',font=font,fill='white')
im.save(r/'output/逐段与转场核验.jpg',quality=92)
blacks=[]
for t in [d['total']-3,d['total']-1.5,d['total']-.1,d['total']-1/30]:
b=subprocess.check_output(['ffmpeg','-v','error','-ss',str(t),'-i',str(out),'-frames:v','1','-pix_fmt','rgb24','-f','rawvideo','-']);a=np.frombuffer(b,np.uint8);blacks.append({'time':t,'mean':float(a.mean()),'max':int(a.max())})
report['end_black']=blacks;assert blacks[-1]['mean']<2
raw=subprocess.check_output(['ffmpeg','-v','error','-i',str(out),'-vn','-ac','2','-ar','48000','-f','f32le','-']);a=np.frombuffer(raw,np.float32).reshape(-1,2);report['audio_peak']=float(np.abs(a).max());report['audio_last_100ms_rms']=float(np.sqrt((a[-4800:]**2).mean()));report['one_second_rms']=[float(np.sqrt((a[i:i+48000]**2).mean())) for i in range(0,len(a),48000)];assert report['audio_peak']<1
for s in d['segments']:
s['subtitles_verified']=all(s['start']+q['end']<=s['start']+s['duration'] for q in s['cues'])
report['segment_count']=len(d['segments']);report['voice_segments']=[s['film'] for s in d['segments'] if s['voices']];report['status']='passed technical checks; contact sheet requires visual review'
(r/'output/验收报告.json').write_text(json.dumps(report,ensure_ascii=False,indent=2));print(json.dumps(report,ensure_ascii=False,indent=2))
- 每个片段与每个转场都抽帧:检查主体构图、字幕对应、黑边、片名黑底及转场无双字幕叠印。
- 对白字幕逐句对应实际人声:稀疏对白要检查声音起止是否完整,不能只用字幕截图证明声音存在。
- 尾音使用 check_audio_tail.py 验证;参数与峰值通过不代表全片一直有音乐。
- 检查末尾图像亮度逐步趋近黑场、音频峰值未超限、文件能正常解码。
- 技术检查与抽帧不是完整听音的替代。复听对白与BGM平衡,留意分离残留、混响、呼吸和突兀压低。
- 播放器不要同时开启外挂 SRT 和已经烧录的字幕,否则会显示两份。
最终 v4:98.5 秒、H.264 + AAC、1280×720、30fps、48kHz 双声道;尾音检查通过,13 段和 12 个转场已抽帧核对。此处的验收只说明已执行的技术检查和视觉检查,不把自动转写通过等同于人声毫无失真。
# 为最新成片创建桌面软链接,不覆盖不属于本项目的现有文件
ln -s "$HOME/Desktop/farewell-mashup/output/离别_Farewell.mp4" "$HOME/Desktop/离别_成品.mp4"
成片保留一个稳定文件名,桌面软链接指向它;另存 v2/v3/v4 版本便于回看。移动或删除项目会使软链接失效。版本化配置、原文备份和成片都保留,避免更新后无法解释哪一版对应哪组参数。
九、本次新增的踩坑清单
| 问题 | 表现 | 本次修正 |
|---|---|---|
| 对白规则收得太窄 | 多个镜头只有对白字幕、没有演员声音 | 所有有明确对白字幕的10段补上分离人声 |
| 延迟音乐后未重置PTS | 片尾约3.5秒突然静音,文件时长却正常 | adelay后插入asetpts=N/SR/TB,并增加尾音回归检查 |
| 上下预留黑区 | 主体画面偏小,重复源黑边进一步缩小画面 | 去源黑边,等比覆盖后居中裁切铺满 |
| 用描边代替黑底 | 片名仍与复杂画面直接混在一起 | 纯黑矩形覆盖整个标签区域,白字无描边 |
| 无对白镜头硬加原创文字 | 镜头没有留白,文字不是演员台词 | 删除全部原创旁白,只留画面与音乐 |
| VAD漏检 | Toy Story告别句未进入初次转写 | 目标区间关闭VAD,核对实际词级时间 |
| cropdetect暗场误判 | 把背景暗处当成黑边,过度裁切 | 先抽帧确认,不直接套用自动建议 |
| 文字烧进分段再转场 | 相邻字幕叠印 | 先拼画面,再统一叠文字,窗口避开转场 |
| 片名与字幕共用透明度 | 每次换对白片名也闪烁 | 片名与字幕独立透明事件 |
| 只验文件时长和峰值 | 静音填充的尾部逃过验收 | 增加短窗口RMS、独立人声与最终输出检查 |
| 旧参数残留 | 61/97.5/98.5秒与4/10处人声混用 | 正文、JSON、代码、SRT、报告统一到最终版本 |
| 改镜头未清人声缓存 | 旧人声可能对不上新画面 | 入点/素材/时长变更后清理对应缓存 |
十、素材、署名与下一次复用
音乐署名:“At Rest” Kevin MacLeod (incompetech.com)。Licensed under Creative Commons: By Attribution 4.0 License。许可;作者曲目页面。音乐许可不覆盖电影画面,影片片段来源与音乐应分别记录。
本次素材来源如下。链接只是定位当时使用的公开片段,不意味着不同下载版本、片段使用条件或素材可用性永远不变。
- Casablanca 1942
- Roman Holiday 1953
- Before Sunrise 1995
- Lost in Translation 2003
- Brokeback Mountain 2005
- La La Land 2016
- Interstellar 2014
- Titanic 1997
- Forrest Gump 1994
- Cinema Paradiso 1988
- Good Will Hunting 1997
- E.T. 1982
- Toy Story 3 2010
下一次拿到新主题时,先定情绪线与留白位置,选素材并核对目标对白,再填配置。音乐连续轨、句级分离人声、画面与文字解耦、按帧时间轴、尾音检查这五项保留;片单、入点、对白和曲目可以更换。所有改版都以最终产物重新验收。
FFmpeg 参数参考:adelay、asetpts、atrim。本页的提前截断判断来自本次输入与输出的实际解码对比,不能泛化为所有 FFmpeg 延迟链路必然出错。
