Compare commits

..

2 Commits

Author SHA1 Message Date
auto-approve-bot 05309b170d ci: harden check-frontend-only with git diff fallback
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Successful in 1s
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Check if frontend-only change (pull_request) Successful in 2s
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
CI/CD Pipeline / PR Build API Image (pull_request) Successful in 33s
CI/CD Pipeline / PR Build Worker Image (pull_request) Successful in 34s
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 58s
CI/CD Pipeline / PR Build Web Image (pull_request) Successful in 1m25s
CI/CD Pipeline / Frontend Unit Tests (pull_request) Successful in 1m39s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 1m46s
CI/CD Pipeline / Frontend Lint (pull_request) Successful in 1m47s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 1m50s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 1m52s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 2m47s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 4m44s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 2m7s
ACR Cleanup / ACR Image Cleanup (pull_request_target) Successful in 55s
Preview Cleanup / Cleanup Preview Environment (pull_request) Successful in 1m0s
AI Code Review / AI Code Review (pull_request) Successful in 6m18s
CI/CD Pipeline / Unit Tests (pull_request) Failing after 9m29s
CI/CD Pipeline / Build Production API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Web Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Production (pull_request) Has been skipped
CI/CD Pipeline / CI Gate (pull_request) Failing after 4s
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
CI/CD Pipeline / Canary Release to Production (pull_request) Has been skipped
- 加 pull-requests: read 权限,避免 github.token 调 /pulls/files 403
- check-frontend-only 优先用 git diff 判断 PR 改动范围,API 失败时自动 fallback
- 修复 #48555 Check if frontend-only change 失败导致 unit-tests 跑全量 diff-cover 被误拦

Closes: CI 纯前端 PR 误跑 Python 测试被覆盖率门槛拦截问题
2026-09-09 13:17:30 +08:00
auto-approve-bot 474918ec66 fix: add heading to Step 5 single video for E2E test
CI/CD Pipeline / Check push changed paths (pull_request) Has been skipped
CI/CD Pipeline / Build Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Build Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging API Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Retag skipped Staging Web Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Staging (Watchtower auto-deploy) (pull_request) Has been skipped
CI/CD Pipeline / Staging API Integration Tests (pull_request) Has been skipped
CI/CD Pipeline / Staging E2E Tests (pull_request) Has been skipped
CI/CD Pipeline / ACR Image Cleanup (pull_request) Has been skipped
Preview Deploy / Deploy Preview Environment (pull_request) Successful in 1m57s
PR Automation / Auto Approve on CI Green (pull_request) Successful in 3m5s
AI Code Review / AI Code Review (pull_request) Successful in 6m46s
CI/CD Pipeline / Check if frontend-only change (pull_request) Failing after 11m30s
CI/CD Pipeline / Dedup Check - skip PR tests when covered by push pipeline (pull_request) Failing after 11m42s
CI/CD Pipeline / PR Build API Image (pull_request) Has been skipped
CI/CD Pipeline / PR Build Web Image (pull_request) Has been skipped
CI/CD Pipeline / Frontend Lint (pull_request) Has been skipped
CI/CD Pipeline / PR Build Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Frontend Unit Tests (pull_request) Successful in 1m26s
CI/CD Pipeline / Validate - Python (mypy + alembic) (pull_request) Successful in 1m40s
CI/CD Pipeline / Integration Tests (pull_request) Successful in 1m41s
CI/CD Pipeline / Validate - Style (pull_request) Successful in 2m18s
PR Automation / Auto Merge on CI Green + Approved (pull_request) Successful in 11m29s
CI/CD Pipeline / Validate - Security (pull_request) Successful in 5m9s
CI/CD Pipeline / Unit Tests (pull_request) Failing after 10m4s
CI/CD Pipeline / Build Production API Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Worker Image (pull_request) Has been skipped
CI/CD Pipeline / Build Production Web Image (pull_request) Has been skipped
CI/CD Pipeline / Deploy Production (pull_request) Has been skipped
CI/CD Pipeline / Canary Release to Production (pull_request) Has been skipped
CI/CD Pipeline / CI Gate (pull_request) Failing after 5s
CI/CD Pipeline / Production Browser E2E (pull_request) Has been skipped
Step 5 单视频场景缺少 🎬 确认生成 heading,导致 Playwright E2E 测试
core-generation.spec.ts 超时失败(getByRole('heading', { name: '🎬 确认生成' }))。

修改 GeneratePage.tsx:
- 在单视频 Step 5 容器顶部加 <h2>🎬 确认生成</h2>
- 外层 div flex 布局改为 column + center,保证垂直居中

Closes: E2E 连续 6 次 develop push 失败问题
2026-09-09 12:49:07 +08:00
34 changed files with 309 additions and 1875 deletions
@@ -1,45 +0,0 @@
"""lipsync_jobs 增加 TTS 直生字段(voice_id/script_text/speed/emotion
Revision ID: 073_add_lipsync_tts_fields
Revises: 072_add_ai_avatar_render
Create Date: 2026-09-09
"""
import sqlalchemy as sa
from alembic import op
revision = "073_add_lipsync_tts_fields"
down_revision = "072_add_ai_avatar_render"
branch_labels = None
depends_on = None
def upgrade() -> None:
# 对口型支持「传音色 + 文案直接生成」:后端内部先 TTS 合成音频再提交对口型
op.add_column(
"lipsync_jobs",
sa.Column("voice_id", sa.String(200), nullable=False, server_default=""),
)
op.add_column(
"lipsync_jobs",
sa.Column("script_text", sa.Text(), nullable=False, server_default=""),
)
op.add_column(
"lipsync_jobs",
sa.Column("speed", sa.Float(), nullable=False, server_default=sa.text("1.0")),
)
op.add_column(
"lipsync_jobs",
sa.Column("emotion", sa.String(20), nullable=False, server_default=""),
)
# audio_url 改为可空:直生模式下音频由后端 TTS 合成后回填
op.alter_column("lipsync_jobs", "audio_url", existing_type=sa.Text(), nullable=True)
def downgrade() -> None:
op.alter_column("lipsync_jobs", "audio_url", existing_type=sa.Text(), nullable=False)
op.drop_column("lipsync_jobs", "emotion")
op.drop_column("lipsync_jobs", "speed")
op.drop_column("lipsync_jobs", "script_text")
op.drop_column("lipsync_jobs", "voice_id")
@@ -1,36 +0,0 @@
"""ai_avatar_render_jobs.script_id 放宽为可空串(手动文案直生场景不关联文案库)
Revision ID: 074_render_script_id_optional
Revises: 073_add_lipsync_tts_fields
Create Date: 2026-09-09
"""
import sqlalchemy as sa
from alembic import op
revision = "074_render_script_id_optional"
down_revision = "073_add_lipsync_tts_fields"
branch_labels = None
depends_on = None
def upgrade() -> None:
# 列保持 NOT NULL(空串占位),仅应用层允许不传;这里显式补 server_default 防止历史约束歧义
with op.batch_alter_table("ai_avatar_render_jobs") as batch:
batch.alter_column(
"script_id",
existing_type=sa.String(length=36),
nullable=False,
server_default="",
)
def downgrade() -> None:
with op.batch_alter_table("ai_avatar_render_jobs") as batch:
batch.alter_column(
"script_id",
existing_type=sa.String(length=36),
nullable=False,
server_default=None,
)
+5 -39
View File
@@ -17,10 +17,7 @@ from app.dependencies import get_db_session
from app.schemas.ai_avatar_render import (
AiAvatarRenderJobResponse,
CreateAiAvatarRenderRequest,
SmartCoverRequest,
SmartCoverResponse,
)
from app.services.ai_avatar_cover_service import generate_smart_cover
from app.services.ai_avatar_render_service import (
AiAvatarRenderError,
AiAvatarRenderService,
@@ -52,7 +49,7 @@ def create_render_job(
"""
try:
job = svc.create_render_job(
user_id=current_user.user.id,
user_id=current_user.id,
lipsync_job_id=body.lipsync_job_id,
script_id=body.script_id,
b_roll_segments=[s.model_dump() for s in body.b_roll_segments],
@@ -97,7 +94,7 @@ def list_render_jobs(
):
"""获取 AI 数字人渲染任务列表."""
items, total = svc.list_render_jobs(
user_id=current_user.user.id,
user_id=current_user.id,
project_id=project_id,
status=status,
offset=offset,
@@ -121,7 +118,7 @@ def get_render_job(
svc: AiAvatarRenderService = Depends(_get_service),
):
"""获取渲染任务详情."""
job = svc.get_render_job(job_id, current_user.user.id)
job = svc.get_render_job(job_id, current_user.id)
if job is None:
raise HTTPException(status_code=404, detail="渲染任务不存在")
return job
@@ -137,7 +134,7 @@ def cancel_render_job(
svc: AiAvatarRenderService = Depends(_get_service),
):
"""取消渲染任务(仅 pending 状态可取消)."""
job = svc.cancel_render_job(job_id, current_user.user.id)
job = svc.cancel_render_job(job_id, current_user.id)
if job is None:
raise HTTPException(status_code=404, detail="渲染任务不存在")
if job.status != "cancelled":
@@ -158,7 +155,7 @@ def retry_render_job(
svc: AiAvatarRenderService = Depends(_get_service),
):
"""重试失败的渲染任务."""
job = svc.retry_render_job(job_id, current_user.user.id)
job = svc.retry_render_job(job_id, current_user.id)
if job is None:
raise HTTPException(status_code=404, detail="渲染任务不存在")
if job.status != "pending":
@@ -176,34 +173,3 @@ def retry_render_job(
logger.warning("Celery 任务提交失败,重试任务已重置但未触发执行: %s", job.id)
return job
# ── POST /smart-cover — 智能获取封面(MediaKit 抽帧 + 评分选帧)────────
@router.post("/smart-cover", response_model=SmartCoverResponse)
def generate_avatar_smart_cover(
body: SmartCoverRequest,
current_user: AuthenticatedUser = Depends(get_current_user),
) -> SmartCoverResponse:
"""智能获取数字人视频封面.
复用智能剪辑的 MediaKit 抽帧 + 质量评分选最佳帧逻辑(非 FFmpeg 简单截帧),
并将选中帧转存到自家 OSS,返回非临时的封面公网 URL。
前端「智能获取封面」按钮可直接调用本接口;不依赖渲染任务完成。
"""
video_url = (body.video_url or "").strip()
if not video_url.startswith(("http://", "https://")):
raise HTTPException(status_code=400, detail="video_url 必须是合法的 HTTP/HTTPS URL")
cover_url = generate_smart_cover(video_url, max_frames=body.max_frames)
if not cover_url:
return SmartCoverResponse(
cover_url="",
status="fallback_failed",
message="智能抽帧失败(MediaKit 不可用或抽帧异常),请稍后重试",
)
logger.info("智能封面生成成功: user=%s", current_user.user.id)
return SmartCoverResponse(cover_url=cover_url, status="completed")
+44 -46
View File
@@ -13,15 +13,11 @@ from __future__ import annotations
import logging
from app.auth import AuthenticatedUser, get_current_user
from app.dependencies import (
get_cosyvoice_service,
get_db_session,
get_voice_clone_profile_repository,
)
from app.dependencies import get_cosyvoice_service, get_db_session, get_voice_clone_profile_repository
from app.schemas.lipsync import CreateLipsyncJobRequest, LipsyncJobResponse
from app.services.lipsync_service import LipsyncService
from app.services.mediakit_client import MediaKitError
from fastapi import APIRouter, BackgroundTasks, Depends, HTTPException, Query
from fastapi import APIRouter, Depends, HTTPException, Query
from sqlalchemy.orm import Session
from packages.application.cosyvoice_service import CosyVoiceError, CosyVoiceService
@@ -33,16 +29,35 @@ router = APIRouter()
def _get_service(
db: Session = Depends(get_db_session),
voice_clone_repo=Depends(get_voice_clone_profile_repository),
cosyvoice_service: CosyVoiceService = Depends(get_cosyvoice_service),
) -> LipsyncService:
# voice_clone_repo 用于克隆音色 profile 解析;cosyvoice_service 用于 TTS 直生
# (TTS 合成、音色解析、错误码归一化都在 LipsyncService 内部完成)
return LipsyncService(
db,
cosyvoice_service=cosyvoice_service,
voice_clone_repo=voice_clone_repo,
)
return LipsyncService(db, cosyvoice_service=cosyvoice_service)
def _resolve_voice_id(
raw_voice_id: str,
user_id: str,
voice_clone_repo,
) -> str:
"""解析 voice_id:支持预设音色 ID 或克隆音色 profile UUID.
与 TTS 路由保持一致:命中 profile → 校验归属 → 取 CosyVoice voice_id。
"""
try:
profile = voice_clone_repo.get(raw_voice_id)
except Exception as exc:
logger.error("查询克隆音色失败: voice_id=%s, error=%s", raw_voice_id, exc)
raise HTTPException(
status_code=400,
detail=f"voice_id 无效: {raw_voice_id}",
) from exc
if profile is not None:
if profile.user_id != user_id:
raise HTTPException(status_code=403, detail="无权访问该音色")
if not profile.voice_id:
raise HTTPException(status_code=400, detail="音色克隆尚未完成,请稍后再试")
return profile.voice_id
return raw_voice_id
# ── POST /jobs — 提交对口型任务 ───────────────────────────────────────────
@@ -53,22 +68,22 @@ def create_lipsync_job(
body: CreateLipsyncJobRequest,
current_user: AuthenticatedUser = Depends(get_current_user),
svc: LipsyncService = Depends(_get_service),
voice_clone_repo=Depends(get_voice_clone_profile_repository),
):
"""提交对口型任务.
#1809/#1822: 前端传 {video_url, voice_id, script_text, speed?, emotion?}
后端内部解析音色、调 TTS 合成音频、转存 OSS,再提交 MediaKit
也支持直接传 {video_url, audio_url}。
#1809: 前端传 {voice_id, script_text, video_url}
后端内部调 TTS 合成音频,再提交 MediaKit
"""
# 解析 voice_id(支持克隆音色 profile UUID
actual_voice_id = _resolve_voice_id(body.voice_id, current_user.id, voice_clone_repo)
try:
job = svc.create_job(
user_id=current_user.user.id,
user_id=current_user.id,
video_url=body.video_url,
audio_url=body.audio_url,
voice_id=body.voice_id,
voice_id=actual_voice_id,
script_text=body.script_text,
speed=body.speed,
emotion=body.emotion,
enable_video_loop=body.enable_video_loop,
project_id=body.project_id,
)
@@ -82,20 +97,12 @@ def create_lipsync_job(
detail={"code": "TTSSynthesisFailed", "message": str(exc)},
) from exc
except MediaKitError as exc:
# TTS 合成失败 / 音色无权访问 → 400/403MediaKit 提交失败 → 502
status_code = 502
if exc.code in ("VoiceForbidden",):
status_code = 403
elif exc.code in ("InvalidInput", "TTSInvalidParam", "VoiceNotReady"):
status_code = 400
elif exc.code == "TTSSynthesisFailed":
status_code = 502
raise HTTPException(
status_code=status_code,
status_code=502,
detail={
"code": exc.code,
"message": str(exc),
"request_id": getattr(exc, "request_id", ""),
"request_id": exc.request_id,
},
) from exc
except Exception as exc:
@@ -123,7 +130,7 @@ def list_lipsync_jobs(
):
"""获取对口型任务列表."""
items, total = svc.list_jobs(
user_id=current_user.user.id,
user_id=current_user.id,
project_id=project_id,
status=status,
offset=offset,
@@ -143,22 +150,13 @@ def list_lipsync_jobs(
@router.get("/jobs/{job_id}", response_model=LipsyncJobResponse)
def get_lipsync_job(
job_id: str,
background: BackgroundTasks,
current_user: AuthenticatedUser = Depends(get_current_user),
svc: LipsyncService = Depends(_get_service),
):
"""获取对口型任务详情.
非终态任务:先返回 DB 缓存,挂后台刷新(下次轮询拿到新状态),
避免 MediaKit 慢响应阻塞前端轮询。
"""
job = svc.get_job(job_id, current_user.user.id)
"""获取对口型任务详情."""
job = svc.get_job(job_id, current_user.id)
if job is None:
raise HTTPException(status_code=404, detail="任务不存在")
if job.status not in ("completed", "failed"):
background.add_task(svc.refresh_job_status, job_id, current_user.user.id)
return job
@@ -172,7 +170,7 @@ def refresh_lipsync_job(
svc: LipsyncService = Depends(_get_service),
):
"""从 MediaKit 拉取最新状态并更新."""
job = svc.refresh_job_status(job_id, current_user.user.id)
job = svc.refresh_job_status(job_id, current_user.id)
if job is None:
raise HTTPException(status_code=404, detail="任务不存在")
return job
@@ -188,7 +186,7 @@ def cancel_lipsync_job(
svc: LipsyncService = Depends(_get_service),
):
"""取消对口型任务(仅 pending/submitted 状态可取消)."""
job = svc.cancel_job(job_id, current_user.user.id)
job = svc.cancel_job(job_id, current_user.id)
if job is None:
raise HTTPException(status_code=404, detail="任务不存在")
if job.status != "cancelled":
+1 -10
View File
@@ -173,14 +173,6 @@ def synthesize(
# job.voice_id 统一存解析后的 CosyVoice voice_id
actual_voice_id = resolved_profile.voice_id
# 语速/情绪等合成参数随 metadata 落库,workflow 提交 CosyVoice 时读取透传
synthesis_meta = {
"speed": request.speed,
"emotion": request.emotion or "",
}
if request.metadata_:
synthesis_meta.update(request.metadata_)
use_case = CreateTTSJobUseCase(repository)
job = use_case.execute(
user_id=user_id,
@@ -188,7 +180,7 @@ def synthesize(
voice_id=actual_voice_id,
voice_model=request.voice_model,
voice_clone_profile_id=voice_clone_profile_id,
metadata=synthesis_meta,
metadata=request.metadata_,
)
# 提交 CosyVoice 合成任务
@@ -575,7 +567,6 @@ def preview_tts(
text=request.text,
voice_id=actual_voice_id,
speed=request.speed,
emotion=request.emotion,
)
except CosyVoiceError as e:
raise HTTPException(
+6 -19
View File
@@ -50,7 +50,7 @@ class CreateAiAvatarRenderRequest(BaseModel):
"""创建渲染任务请求."""
lipsync_job_id: str = Field(..., description="对口型任务 ID")
script_id: str = Field("", description="文案 ID(选自文案库时传;手动输入文案直生场景可留空)")
script_id: str = Field(..., description="文案 ID")
b_roll_segments: list[BRollSegment] = Field(default_factory=list, description="B-roll 片段列表")
title_config: dict[str, Any] = Field(default_factory=dict, description="标题配置")
cover_config: dict[str, Any] = Field(default_factory=dict, description="封面配置")
@@ -67,8 +67,10 @@ class CreateAiAvatarRenderRequest(BaseModel):
@field_validator("script_id")
@classmethod
def validate_script_id(cls, v: str) -> str:
# script_id 可选:手动输入文案(TTS 直生)场景不关联文案库条目
return (v or "").strip()
v = v.strip()
if not v:
raise ValueError("script_id 不能为空")
return v
class AiAvatarRenderJobResponse(BaseModel):
@@ -78,7 +80,7 @@ class AiAvatarRenderJobResponse(BaseModel):
user_id: str
project_id: str
lipsync_job_id: str
script_id: str = ""
script_id: str
b_roll_segments: list[dict[str, Any]]
title_config: dict[str, Any]
cover_config: dict[str, Any]
@@ -107,18 +109,3 @@ class AiAvatarRenderProgressResponse(BaseModel):
output_cover_url: str
output_duration: float
error_message: str
class SmartCoverRequest(BaseModel):
"""智能封面请求 — MediaKit 抽帧 + 质量评分选最佳帧."""
video_url: str = Field(..., description="数字人视频 URL(对口型/渲染成片)")
max_frames: int = Field(5, ge=1, le=10, description="抽帧数量(默认 5")
class SmartCoverResponse(BaseModel):
"""智能封面响应."""
cover_url: str = Field("", description="封面图公网 URL(OSS,非临时);失败为空")
status: str = Field("completed", description="completed / fallback_failed")
message: str = Field("", description="失败原因(如有)")
+30 -53
View File
@@ -1,17 +1,11 @@
"""对口型 API Schema 定义 — #1796 / #1809 / #1822.
支持两种输入模式(二选一):
1. TTS 直生模式(推荐):传 voice_id + script_text+ speed/emotion),
后端内部先调 CosyVoice 合成音频,再提交 MediaKit 对口型。
2. 直接音频模式:传 video_url + audio_url(音频已由调用方准备好)。
"""
"""对口型 API Schema 定义 — #1796, #1809 参数调整."""
from __future__ import annotations
from datetime import datetime
from typing import Optional
from pydantic import BaseModel, Field, model_validator
from pydantic import BaseModel, Field, field_validator
class LipsyncJobResponse(BaseModel):
@@ -23,10 +17,6 @@ class LipsyncJobResponse(BaseModel):
video_url: str
audio_url: str
enable_video_loop: bool
voice_id: str = ""
script_text: str = ""
speed: float = 1.0
emotion: str = ""
mediakit_task_id: str
status: str
output_video_url: str
@@ -43,58 +33,45 @@ class LipsyncJobResponse(BaseModel):
class CreateLipsyncJobRequest(BaseModel):
"""创建对口型任务请求.
"""创建对口型任务请求 — #1809.
两种模式(二选一):
- TTS 直生:voice_id + script_text 必填(+ 可选 speed/emotion);audio_url 留空
- 直接音频:video_url + audio_url 必填。
前端传 {voice_id, script_text, video_url}
后端内部调 TTS 生成 audio_url 再提交 MediaKit
"""
video_url: str = Field(..., description="人物视频 URL(MP4,≤30min,单人真人)")
# 模式 2:直接音频
audio_url: str = Field("", description="驱动音频 URLmp3/aac/wav/m4a/flac);直生模式留空")
# 模式 1TTS 直生
voice_id: str = Field("", description="音色 ID(预置音色或克隆音色 profile UUID")
script_text: str = Field("", description="要合成的文案(直生模式必填,最长 5000 字符)")
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速(0.5-2.0),默认 1.0")
emotion: str = Field("", description="情绪(natural/excited/calm/friendly 或中文 自然/兴奋/沉稳/亲切)")
voice_id: str = Field(..., description="音色 ID(预设音色或克隆音色 profile ID)")
script_text: str = Field(..., description="要合成的脚本文本")
enable_video_loop: bool = Field(False, description="音频长于视频时是否循环画面")
project_id: str = Field("", description="项目 ID(可选)")
@model_validator(mode="after")
def _validate_input_mode(self) -> "CreateLipsyncJobRequest":
video = (self.video_url or "").strip()
if not video:
@field_validator("video_url")
@classmethod
def validate_video_url(cls, v: str) -> str:
v = v.strip()
if not v:
raise ValueError("video_url 不能为空")
if not video.startswith(("http://", "https://")):
if not v.startswith(("http://", "https://")):
raise ValueError("video_url 必须是 HTTP/HTTPS URL")
lower = video.lower().split("?")[0]
lower = v.lower().split("?")[0]
if not lower.endswith(".mp4"):
raise ValueError("video_url 仅支持 MP4 格式")
return v
has_audio = bool((self.audio_url or "").strip())
has_tts = bool((self.voice_id or "").strip()) and bool((self.script_text or "").strip())
@field_validator("voice_id")
@classmethod
def validate_voice_id(cls, v: str) -> str:
v = v.strip()
if not v:
raise ValueError("voice_id 不能为空")
return v
if not has_audio and not has_tts:
raise ValueError(
"必须提供驱动音频:要么传 audio_url(直接音频模式),"
"要么同时传 voice_id + script_textTTS 直生模式)"
)
if has_tts and len(self.script_text) > 5000:
@field_validator("script_text")
@classmethod
def validate_script_text(cls, v: str) -> str:
v = v.strip()
if not v:
raise ValueError("script_text 不能为空")
if len(v) > 5000:
raise ValueError("script_text 最长 5000 字符")
if has_audio:
au = self.audio_url.strip()
if not au.startswith(("http://", "https://")):
raise ValueError("audio_url 必须是 HTTP/HTTPS URL")
au_lower = au.lower().split("?")[0]
allowed = (".mp3", ".aac", ".wav", ".m4a", ".flac")
if not any(au_lower.endswith(ext) for ext in allowed):
raise ValueError(f"audio_url 格式不支持,仅支持: {', '.join(allowed)}")
self.audio_url = au
return self
return v
-2
View File
@@ -16,7 +16,6 @@ class TTSSynthesizeRequest(BaseModel):
output_name: str = Field("", description="输出文件名")
language: str = Field("zh-CN", description="语言")
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速")
emotion: str = Field("", description="情绪(natural/excited/calm/friendly,或中文 自然/兴奋/沉稳/亲切)")
voice_model: str = Field("", description="语音模型名称")
voice_clone_profile_id: str = Field("", description="关联的音色克隆档案 ID")
format: str = Field("mp3", description="输出格式(mp3/wav/pcm")
@@ -110,7 +109,6 @@ class TTSPreviewRequest(BaseModel):
text: str = Field(..., min_length=1, max_length=200, description="合成文本,限制 200 字")
voice_id: str = Field(..., min_length=1, description="音色 ID")
speed: float = Field(1.0, ge=0.5, le=2.0, description="语速")
emotion: str = Field("", description="情绪(natural/excited/calm/friendly,或中文)")
pitch: float = Field(1.0, ge=0.5, le=2.0, description="音调(预留,当前未使用)")
@@ -1,162 +0,0 @@
"""AI 数字人封面服务 — 复用智能剪辑的 MediaKit 抽帧 + 质量评分选最佳帧.
与 generation_cover.py 的智能选帧能力对齐(不再用 FFmpeg 简单截帧):
1. MediaKit extract_frames 抽取多帧(默认 5 帧,SpecifiedFrames 策略)
2. cover_frame_scorer.score_frames 按清晰度/亮度/色彩评分选最佳
3. 下载最佳帧并转存 OSS,返回公网封面 URL
降级:MediaKit 不可用或抽帧失败时返回空字符串,由调用方决定回退策略。
"""
from __future__ import annotations
import logging
import tempfile
import uuid
from pathlib import Path
from typing import Optional
logger = logging.getLogger(__name__)
def select_best_cover_frame(video_url: str, *, max_frames: int = 5) -> str:
"""从视频抽取多帧并评分选最佳帧,返回最佳帧的临时 URL.
Args:
video_url: 可公网访问的视频 URL
max_frames: 抽帧数量
Returns:
最佳帧图片 URL;失败返回空字符串
"""
if not video_url:
return ""
try:
from packages.shared.cover_frame_scorer import score_frames
from packages.shared.mediakit_client import get_mediakit_client
mk = get_mediakit_client()
if not mk.is_available:
logger.warning("[数字人封面] MediaKit 未配置,无法智能抽帧")
return ""
snapshots = mk.extract_frames(
video_url=video_url,
strategy="SpecifiedFrames",
max_frames=max_frames,
poll_interval=2.0,
max_poll_attempts=5,
max_retries=0,
)
if not snapshots:
logger.warning("[数字人封面] MediaKit 未返回帧: %s", video_url[:80])
return ""
if len(snapshots) == 1:
return snapshots[0].get("image_url") or snapshots[0].get("url") or ""
# 下载各帧评分
import httpx
candidates = []
for snap in snapshots:
url = snap.get("image_url") or snap.get("url") or ""
if not url:
continue
tmp_path: Optional[str] = None
try:
resp = httpx.get(url, timeout=15, follow_redirects=True)
resp.raise_for_status()
with tempfile.NamedTemporaryFile(suffix=".jpg", delete=False) as tmp:
tmp.write(resp.content)
tmp_path = tmp.name
candidates.append({"image_path": tmp_path, "url": url})
except Exception:
candidates.append({"image_path": None, "url": url, "score": 0.0})
if not candidates:
return snapshots[0].get("image_url") or snapshots[0].get("url") or ""
scored = score_frames(candidates)
best = scored[0] if scored else None
best_url = best.get("url", "") if best else ""
# 清理临时文件
for c in candidates:
p = c.get("image_path")
if p:
try:
Path(p).unlink(missing_ok=True)
except Exception:
pass
logger.info(
"[数字人封面] 智能选帧完成: candidates=%d best_score=%s",
len(candidates),
best.get("score") if best else "n/a",
)
return best_url
except Exception:
logger.warning("[数字人封面] 智能选帧失败", exc_info=True)
return ""
def persist_cover_to_oss(frame_url: str, *, job_id: str = "", prefix: str = "ai-avatar/covers") -> str:
"""下载帧图并转存到 OSS,返回公网封面 URL.
Args:
frame_url: MediaKit 返回的临时帧图 URL
job_id: 关联任务 ID(用于 OSS key 命名)
prefix: OSS key 前缀
Returns:
OSS 公网 URL;失败回退原始 frame_url
"""
if not frame_url:
return ""
tmp_path: Optional[str] = None
try:
import httpx
resp = httpx.get(frame_url, timeout=30, follow_redirects=True)
resp.raise_for_status()
if not resp.content:
return frame_url
with tempfile.NamedTemporaryFile(suffix=".jpg", delete=False) as tmp:
tmp.write(resp.content)
tmp_path = tmp.name
from packages.shared.storage import get_shared_storage_service
storage = get_shared_storage_service()
token = job_id or uuid.uuid4().hex[:12]
cover_key = f"{prefix}/{token}/cover_{uuid.uuid4().hex[:8]}.jpg"
public_url = storage.upload_file(
file_or_path=tmp_path,
storage_key=cover_key,
content_type="image/jpeg",
)
logger.info("[数字人封面] 封面已转存 OSS: key=%s", cover_key)
return public_url or frame_url
except Exception:
logger.warning("[数字人封面] 封面转存 OSS 失败,返回原始 URL", exc_info=True)
return frame_url
finally:
if tmp_path:
try:
Path(tmp_path).unlink(missing_ok=True)
except Exception:
pass
def generate_smart_cover(video_url: str, *, job_id: str = "", max_frames: int = 5) -> str:
"""一站式:MediaKit 智能抽帧选最佳 → 转存 OSS,返回封面公网 URL.
供独立封面接口与渲染管线复用。失败返回空字符串。
"""
best_frame = select_best_cover_frame(video_url, max_frames=max_frames)
if not best_frame:
return ""
return persist_cover_to_oss(best_frame, job_id=job_id)
@@ -53,8 +53,8 @@ class AiAvatarRenderService:
*,
user_id: str,
lipsync_job_id: str,
script_id: str = "",
b_roll_segments: list[dict[str, Any]] | None = None,
script_id: str,
b_roll_segments: list[dict[str, Any]],
title_config: dict[str, Any],
cover_config: dict[str, Any],
project_id: str = "",
@@ -83,19 +83,17 @@ class AiAvatarRenderService:
if not lipsync_job.output_video_url:
raise AiAvatarRenderError("对口型任务输出视频 URL 为空", code="LipsyncJobNoOutput")
# 2. 验证文案归属(仅当选了文案库条目时;手动输入文案直生场景 script_id 可空)
script_id = (script_id or "").strip()
if script_id:
script = (
self.db.query(ScriptModel)
.filter(
ScriptModel.id == script_id,
ScriptModel.user_id == user_id,
)
.first()
# 2. 验证文案归属
script = (
self.db.query(ScriptModel)
.filter(
ScriptModel.id == script_id,
ScriptModel.user_id == user_id,
)
if script is None:
raise AiAvatarRenderError("文案不存在或无权访问", code="ScriptNotFound")
.first()
)
if script is None:
raise AiAvatarRenderError("文案不存在或无权访问", code="ScriptNotFound")
# 3. 创建渲染任务
job_id = str(uuid.uuid4())
@@ -105,7 +103,7 @@ class AiAvatarRenderService:
project_id=project_id,
lipsync_job_id=lipsync_job_id,
script_id=script_id,
b_roll_segments=[s if isinstance(s, dict) else s.model_dump() for s in (b_roll_segments or [])],
b_roll_segments=[s if isinstance(s, dict) else s.model_dump() for s in b_roll_segments],
title_config=title_config,
cover_config=cover_config,
status="pending",
@@ -290,22 +288,7 @@ class AiAvatarRenderService:
output_video_url = self._upload_to_oss(output_video_path, f"ai-avatar/{job_id}/output.mp4")
job.output_video_url = output_video_url
# 封面:优先复用智能剪辑的 MediaKit 抽帧 + 质量评分选最佳帧;
# MediaKit 不可用时回退到 FFmpeg 已按 cover_config 抽取的 cover_path
smart_cover_url = ""
if output_video_url:
try:
from app.services.ai_avatar_cover_service import (
generate_smart_cover,
)
smart_cover_url = generate_smart_cover(output_video_url, job_id=job_id, max_frames=5)
except Exception:
logger.warning("智能封面(MediaKit)失败,回退 FFmpeg 封面 job_id=%s", job_id, exc_info=True)
if smart_cover_url:
job.output_cover_url = smart_cover_url
elif cover_path:
if cover_path:
output_cover_url = self._upload_to_oss(cover_path, f"ai-avatar/{job_id}/cover.jpg")
job.output_cover_url = output_cover_url
+43 -153
View File
@@ -1,16 +1,15 @@
"""对口型 Service — #1796 MediaKit 对口型业务逻辑, #1809 参数调整.
职责:
- 创建/查询对口型任务
- 双输入模式:TTS 直生(voice_id + script_text,内部先合成音频转存 OSS)或直接音频(audio_url
- 创建/查询/取消对口型任务
- 调用 TTS 合成音频(#1809:前端不再传 audio_url
- 调用 MediaKit 客户端提交异步任务
- 轮询更新任务状态(中间状态同步 DB,成片转存自家 OSS)
- 轮询更新任务状态
- 用户隔离(每个用户只能操作自己的任务)
"""
from __future__ import annotations
import io
import logging
import uuid
from datetime import datetime, timezone
@@ -27,9 +26,7 @@ from app.services.mediakit_client import (
from sqlalchemy.orm import Session
from packages.adapters.sqlalchemy_impl.models import LipsyncJobModel
from packages.application.cosyvoice_service import CosyVoiceError, normalize_emotion
from packages.shared.storage import get_shared_storage_service
from packages.shared.url_security import ALLOWED_AUDIO_MIME_TYPES, safe_download_bytes
from packages.application.cosyvoice_service import CosyVoiceError, CosyVoiceService
logger = logging.getLogger(__name__)
@@ -41,98 +38,19 @@ class LipsyncService:
self,
db: Session,
client: Optional[MediaKitClient] = None,
cosyvoice_service=None,
voice_clone_repo=None,
cosyvoice_service: Optional[CosyVoiceService] = None,
):
self.db = db
self.client = client or get_mediakit_client()
self._cosyvoice = cosyvoice_service
self._voice_clone_repo = voice_clone_repo
self._cosyvoice_service = cosyvoice_service
def _get_cosyvoice(self):
"""延迟获取 CosyVoiceService(与 tts 路由一致,含 OSS 预签名配置)."""
if self._cosyvoice is None:
@property
def cosyvoice_service(self) -> CosyVoiceService:
if self._cosyvoice_service is None:
from app.dependencies import get_cosyvoice_service
self._cosyvoice = get_cosyvoice_service()
return self._cosyvoice
def _resolve_voice_id(self, voice_id: str, user_id: str) -> str:
"""将克隆音色 profile UUID 解析为 CosyVoice voice_id。
与 /tts/synthesize 保持一致:命中 profile → 校验归属 → 返回其 voice_id;
未命中(预置音色 ID 或克隆 CosyVoice voice_id)原样返回。
"""
if not voice_id:
return ""
if self._voice_clone_repo is None:
try:
from app.dependencies import get_voice_clone_profile_repository
self._voice_clone_repo = get_voice_clone_profile_repository(self.db)
except Exception:
return voice_id
try:
profile = self._voice_clone_repo.get(voice_id)
except Exception:
return voice_id
if profile is None:
return voice_id
if getattr(profile, "user_id", "") != user_id:
raise MediaKitError("无权访问该音色", code="VoiceForbidden")
if not getattr(profile, "voice_id", ""):
raise MediaKitError("音色克隆尚未完成,请稍后再试", code="VoiceNotReady")
return profile.voice_id
def _synthesize_and_persist_audio(
self,
*,
user_id: str,
job_id: str,
voice_id: str,
script_text: str,
speed: float,
emotion: str,
) -> str:
"""TTS 直生:调 CosyVoice 合成音频并转存 OSS,返回可公网访问的音频 URL.
Raises:
MediaKitError: 合成失败
"""
actual_voice_id = self._resolve_voice_id(voice_id, user_id)
cosyvoice = self._get_cosyvoice()
try:
result = cosyvoice.submit_synthesize_task(
text=script_text,
voice_id=actual_voice_id,
speed=speed,
emotion=normalize_emotion(emotion),
)
except CosyVoiceError as exc:
raise MediaKitError(f"TTS 合成失败: {exc}", code="TTSSynthesisFailed") from exc
except ValueError as exc:
raise MediaKitError(f"TTS 参数错误: {exc}", code="TTSInvalidParam") from exc
temp_url = result.get("audio_url", "")
if not temp_url:
raise MediaKitError("TTS 未返回音频 URL", code="TTSNoAudio")
# 转存到自家 OSS,避免临时 URL 过期导致 MediaKit 拉取失败
try:
audio_data = safe_download_bytes(
temp_url,
purpose="lipsync_tts_audio",
allowed_mime_types=ALLOWED_AUDIO_MIME_TYPES,
timeout=60.0,
)
storage = get_shared_storage_service()
storage_key = f"lipsync-tts/{user_id}/{job_id}.mp3"
permanent_url = storage.upload_file(io.BytesIO(audio_data), storage_key, content_type="audio/mpeg")
logger.info("对口型 TTS 音频已转存 OSS: job_id=%s key=%s", job_id, storage_key)
return permanent_url
except Exception as exc:
logger.warning("TTS 音频转存 OSS 失败,回退临时 URL: job_id=%s err=%s", job_id, exc)
return temp_url
self._cosyvoice_service = get_cosyvoice_service()
return self._cosyvoice_service
# ── 创建任务 ──────────────────────────────────────────────────────────
@@ -141,42 +59,47 @@ class LipsyncService:
*,
user_id: str,
video_url: str,
audio_url: str = "",
voice_id: str = "",
script_text: str = "",
speed: float = 1.0,
emotion: str = "",
voice_id: str,
script_text: str,
enable_video_loop: bool = False,
project_id: str = "",
) -> LipsyncJobModel:
"""创建对口型任务并提交到 MediaKit.
两种输入模式:
- TTS 直生:voice_id + script_textaudio_url 留空),后端先合成音频
- 直接音频:提供 audio_url
#1809: 内部调 TTS 合成音频,不再由前端传 audio_url。
Raises:
MediaKitError: TTS 合成或 MediaKit 提交失败
CosyVoiceError: TTS 合成失败
MediaKitError: API 调用失败
"""
# 0. TTS 直生模式:先合成音频(在创建 DB 记录之前完成,失败直接抛出)
if not audio_url:
if not (voice_id and script_text):
raise MediaKitError(
"必须提供 audio_url 或 voice_id+script_text",
code="InvalidInput",
)
# 预合成:用临时 job_id 命名 OSS 对象
pre_job_id = str(uuid.uuid4())
audio_url = self._synthesize_and_persist_audio(
user_id=user_id,
job_id=pre_job_id,
# 1. 调 TTS 合成音频
try:
tts_result = self.cosyvoice_service.synthesize_speech(
text=script_text,
voice_id=voice_id,
script_text=script_text,
speed=speed,
emotion=emotion,
)
audio_url = tts_result.audio_url
except CosyVoiceError as exc:
logger.error("TTS 合成失败: voice_id=%s, error=%s", voice_id, exc)
# 创建失败记录
job_id = str(uuid.uuid4())
job = LipsyncJobModel(
id=job_id,
user_id=user_id,
project_id=project_id,
video_url=video_url,
audio_url="",
enable_video_loop=enable_video_loop,
status="failed",
error_message=f"TTS 合成失败: {exc}",
error_code="TTSSynthesisFailed",
)
self.db.add(job)
self.db.commit()
self.db.refresh(job)
raise
# 1. 创建数据库记录
# 2. 创建数据库记录
job_id = str(uuid.uuid4())
job = LipsyncJobModel(
id=job_id,
@@ -185,10 +108,6 @@ class LipsyncService:
video_url=video_url,
audio_url=audio_url,
enable_video_loop=enable_video_loop,
voice_id=voice_id or "",
script_text=script_text or "",
speed=speed,
emotion=normalize_emotion(emotion),
status="pending",
)
self.db.add(job)
@@ -273,14 +192,11 @@ class LipsyncService:
return job
mk_status = status_data.get("status", STATUS_RUNNING)
logger.info("MediaKit 对口型状态 [%s]: %s", job_id, mk_status)
if mk_status == STATUS_COMPLETED:
result = status_data.get("result", {})
job.status = STATUS_COMPLETED
output_url = result.get("video_url", "")
# MediaKit 输出为临时 URL,转存自家 OSS 防止过期(失败则回退临时 URL)
job.output_video_url = self._persist_output_video(output_url, job_id, user_id)
job.output_video_url = result.get("video_url", "")
job.output_duration = result.get("duration", 0.0)
job.completed_at = datetime.now(timezone.utc)
elif mk_status == STATUS_FAILED:
@@ -289,38 +205,12 @@ class LipsyncService:
job.error_message = error.get("message", "任务执行失败")
job.error_code = error.get("code", "TaskFailed")
job.completed_at = datetime.now(timezone.utc)
else:
# 中间状态(running/processing/queued 等)同步到 DB,避免前端永远卡在 submitted
if isinstance(mk_status, str) and mk_status:
job.status = mk_status
# running 状态只更新时间戳
job.updated_at = datetime.now(timezone.utc)
self.db.commit()
self.db.refresh(job)
return job
def _persist_output_video(self, temp_url: str, job_id: str, user_id: str) -> str:
"""将 MediaKit 输出的临时视频 URL 转存到自家 OSS.
失败时回退返回原始临时 URL,不影响任务完成。
"""
if not temp_url:
return ""
try:
import httpx
with httpx.Client(timeout=180.0, follow_redirects=True) as client:
resp = client.get(temp_url)
resp.raise_for_status()
data = resp.content
storage = get_shared_storage_service()
storage_key = f"lipsync-outputs/{user_id}/{job_id}.mp4"
permanent_url = storage.upload_file(io.BytesIO(data), storage_key, content_type="video/mp4")
logger.info("对口型输出视频已转存 OSS: job_id=%s key=%s", job_id, storage_key)
return permanent_url or temp_url
except Exception as exc:
logger.warning("对口型输出视频转存 OSS 失败,回退临时 URL: job_id=%s err=%s", job_id, exc)
return temp_url
# ── 取消任务 ──────────────────────────────────────────────────────────
def cancel_job(self, job_id: str, user_id: str) -> Optional[LipsyncJobModel]:
-1
View File
@@ -103,7 +103,6 @@ export interface TTSPreviewRequest {
voice_id: string
speed?: number
pitch?: number
emotion?: string // 情绪参数:natural/excited/calm/friendly
}
/** TTS 试听响应 */
-26
View File
@@ -1150,29 +1150,3 @@
flex-direction: column;
gap: 4px;
}
/* ─ 对口型生成弹窗 Spinner ── */
.aa-lipsync-spinner {
width: 48px;
height: 48px;
border: 4px solid #f0f0f5;
border-top-color: #6366f1;
border-radius: 50%;
animation: aa-spin 0.8s linear infinite;
}
@keyframes aa-spin {
to {
transform: rotate(360deg);
}
}
.aa-btn--danger {
background: #ff4d4f;
color: #fff;
border: none;
}
.aa-btn--danger:hover {
background: #ff7875;
}
+16 -181
View File
@@ -19,13 +19,7 @@ import {
createLipsyncJob,
getLipsyncJob,
submitRender,
generateSmartCover,
} from "./api/aiAvatar"
import {
normalizeEmotion,
buildTitleConfigPayload,
buildCoverConfigPayload,
} from "./utils/contract"
/** 面板折叠状态 */
type PanelKey = "video" | "voice" | "script" | "title" | "cover"
@@ -40,15 +34,6 @@ const AiAvatarPage: React.FC = () => {
cover: false,
})
/* ── 对口型生成弹窗 ── */
const [showLipsyncModal, setShowLipsyncModal] = useState(false)
const [lipsyncStatus, setLipsyncStatus] = useState<"generating" | "completed" | "failed">(
"generating",
)
const [lipsyncErrorMessage, setLipsyncErrorMessage] = useState("")
/* ── 智能封面加载态 ── */
const [smartCoverLoading, setSmartCoverLoading] = useState(false)
/* ── 对口型轮询 ── */
const lipsyncTimerRef = useRef<ReturnType<typeof setInterval> | null>(null)
@@ -71,91 +56,46 @@ const AiAvatarPage: React.FC = () => {
return
}
try {
// 显示生成弹窗
setShowLipsyncModal(true)
setLipsyncStatus("generating")
setLipsyncErrorMessage("")
// ① 先按素材 id 拿 file_url(#1809 补充:对齐后端新参数 video_url)
console.log("[对口型] 开始生成:", {
videoId: video.id,
voiceId: voice.voice_id,
voiceType: voice.type,
textLen: state.scriptText.length,
})
const asset = await getAssetById(video.id)
console.log("[对口型] getAssetById 响应:", {
id: asset?.id,
file_url: asset?.file_url?.substring(0, 100),
})
const videoUrl = asset?.file_url
if (!videoUrl) {
console.error("[对口型] file_url 为空,asset:", asset)
setShowLipsyncModal(false)
message.error("获取出镜视频播放地址失败,请重新选择素材")
return
}
// ② 模式A TTS直生:video_url + voice_id + script_text,语速/情绪英文枚举透传(#1822)
const payload = {
// ② voice_id(预设/克隆 UUID 均由后端内部调 TTS+ script_text + video_url
const job = await createLipsyncJob({
voice_id: voice.voice_id,
script_text: state.scriptText,
video_url: videoUrl,
speed: state.speed, // 语速 0.5~2.0
emotion: normalizeEmotion(state.emotion), // natural/excited/calm/friendly
}
console.log("[对口型] createLipsyncJob 请求:", payload)
const job = await createLipsyncJob(payload)
console.log("[对口型] createLipsyncJob 响应:", { id: job.id, status: job.status })
})
state.setLipsyncJob(job)
message.success("对口型任务已提交,生成中…")
// 开始轮询
if (lipsyncTimerRef.current) clearInterval(lipsyncTimerRef.current)
lipsyncTimerRef.current = setInterval(async () => {
try {
const updated = await getLipsyncJob(job.id)
state.setLipsyncJob(updated)
console.log("[对口型] 轮询状态:", {
id: updated.id,
status: updated.status,
error: updated.error_message,
})
if (updated.status === "completed") {
if (updated.status === "completed" || updated.status === "failed") {
if (lipsyncTimerRef.current) clearInterval(lipsyncTimerRef.current)
setLipsyncStatus("completed")
setTimeout(() => {
setShowLipsyncModal(false)
if (updated.status === "completed") {
message.success("对口型视频生成完成")
}, 1000)
} else if (updated.status === "failed") {
if (lipsyncTimerRef.current) clearInterval(lipsyncTimerRef.current)
setLipsyncStatus("failed")
setLipsyncErrorMessage(updated.error_message || "对口型生成失败")
} else {
message.error(updated.error_message || "对口型生成失败")
}
}
} catch (err) {
console.error("[对口型] 轮询错误:", err)
} catch {
// 忽略轮询错误(轮询期间不打扰用户)
}
}, 3000)
} catch (err) {
console.error("[对口型] 创建失败:", {
status: (err as { response?: { status?: number } })?.response?.status,
data: (err as { response?: { data?: unknown } })?.response?.data,
message: err instanceof Error ? err.message : String(err),
})
setShowLipsyncModal(false)
// ② 接口失败弹错误提示,不只 console
console.error("对口型任务创建失败:", err)
message.error(err instanceof Error ? err.message : "对口型任务提交失败,请重试")
}
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [state.selectedVideo, state.selectedVoice, state.scriptText, state.speed, state.emotion])
// 取消对口型生成
const handleCancelLipsync = useCallback(() => {
if (lipsyncTimerRef.current) {
clearInterval(lipsyncTimerRef.current)
lipsyncTimerRef.current = null
}
setShowLipsyncModal(false)
setLipsyncStatus("generating")
setLipsyncErrorMessage("")
}, [])
}, [state.selectedVideo, state.selectedVoice, state.scriptText])
// 清理轮询
useEffect(() => {
@@ -177,10 +117,8 @@ const AiAvatarPage: React.FC = () => {
lipsync_job_id: state.lipsyncJob.id,
script_id: state.script?.id,
b_roll_segments: state.bRollSegments as never,
// 字段映射:build_title_drawtext_filter 真实口径 text/font_size/font_color/position/...
title_config: buildTitleConfigPayload(state.titleConfig),
// cover_config:智能封面 cover_url + 截帧 timestamp
cover_config: buildCoverConfigPayload(state.coverConfig, state.coverConfig.smart_cover_url),
title_config: state.titleConfig as unknown as Record<string, unknown>,
cover_config: state.coverConfig as unknown as Record<string, unknown>,
resolution: state.resolution,
})
message.success("渲染任务已提交,可在视频管理中查看进度")
@@ -200,37 +138,6 @@ const AiAvatarPage: React.FC = () => {
state.resolution,
])
/* ── 智能封面:调后端 MediaKit 选帧接口(#1822 ── */
const handleSmartCover = useCallback(async () => {
// 基于对口型成片抽帧,必须先完成对口型
const videoUrl = state.lipsyncJob?.output_video_url
if (state.lipsyncJob?.status !== "completed" || !videoUrl) {
message.warning("请先生成对口型视频,完成后再智能获取封面")
return
}
setSmartCoverLoading(true)
try {
const res = await generateSmartCover(videoUrl, 5)
if (res.cover_url) {
state.setCoverConfig((prev) => ({
...prev,
mode: "auto_frame",
smart_cover_url: res.cover_url,
thumbnail_url: res.cover_url,
}))
message.success("智能封面已生成")
} else {
message.error(res.message || "智能封面生成失败,请稍后重试")
}
} catch (err) {
console.error("智能封面生成失败:", err)
message.error(err instanceof Error ? err.message : "智能封面生成失败,请重试")
} finally {
setSmartCoverLoading(false)
}
// eslint-disable-next-line react-hooks/exhaustive-deps
}, [state.lipsyncJob])
/* ── 配置汇总 ── */
const summary = {
videoName: state.selectedVideo?.name || null,
@@ -259,7 +166,6 @@ const AiAvatarPage: React.FC = () => {
selectedVideo={state.selectedVideo}
onSelectVideo={() => state.setShowAssetPicker(true)}
onRemoveVideo={state.removeVideo}
titleConfig={state.titleConfig}
/>
</div>
</div>
@@ -329,9 +235,6 @@ const AiAvatarPage: React.FC = () => {
onCoverConfigChange={(partial) =>
state.setCoverConfig((prev) => ({ ...prev, ...partial }))
}
onSmartCover={handleSmartCover}
smartCoverLoading={smartCoverLoading}
canSmartCover={state.lipsyncJob?.status === "completed"}
resolution={state.resolution}
onResolutionChange={state.setResolution}
isGenerating={state.isGenerating}
@@ -373,74 +276,6 @@ const AiAvatarPage: React.FC = () => {
onRemove={state.removeBRollSegment}
/>
)}
{/* 对口型生成弹窗 */}
{showLipsyncModal && (
<div className="aa-modal-overlay">
<div className="aa-modal" onClick={(e) => e.stopPropagation()}>
<div className="aa-modal__header">
<span className="aa-modal__title"></span>
<button className="aa-modal__close" onClick={handleCancelLipsync}>
</button>
</div>
<div
className="aa-modal__body"
style={{
display: "flex",
flexDirection: "column",
alignItems: "center",
padding: "40px 20px",
}}
>
{lipsyncStatus === "generating" && (
<>
<div className="aa-lipsync-spinner" />
<div style={{ marginTop: 20, fontSize: 15, color: "#1a1a2e" }}>
</div>
<div style={{ marginTop: 8, fontSize: 13, color: "#8c8ca1" }}>
</div>
</>
)}
{lipsyncStatus === "completed" && (
<>
<div style={{ fontSize: 48 }}></div>
<div style={{ marginTop: 16, fontSize: 15, color: "#1a1a2e" }}>
</div>
</>
)}
{lipsyncStatus === "failed" && (
<>
<div style={{ fontSize: 48 }}></div>
<div style={{ marginTop: 16, fontSize: 15, color: "#1a1a2e" }}>
</div>
{lipsyncErrorMessage && (
<div style={{ marginTop: 8, fontSize: 13, color: "#ff4d4f" }}>
{lipsyncErrorMessage}
</div>
)}
</>
)}
</div>
<div className="aa-modal__footer">
{lipsyncStatus === "generating" && (
<button className="aa-btn aa-btn--danger" onClick={handleCancelLipsync}>
</button>
)}
{lipsyncStatus !== "generating" && (
<button className="aa-btn" onClick={handleCancelLipsync}>
</button>
)}
</div>
</div>
</div>
)}
</div>
)
}
+4 -25
View File
@@ -1,5 +1,5 @@
/**
* AI数字人 — API 封装#1822 契约对齐)
* AI数字人 — API 封装
*/
import apiClient from "@/api/client"
import type { Script, LipsyncJob, RenderJob, BRollSegment } from "../types"
@@ -28,26 +28,17 @@ export const deleteScript = async (id: string): Promise<void> => {
await apiClient.delete(`/scripts/${id}`)
}
/* ── 素材单查(拿到 file_url 作为对口型的 video_url ── */
/* ── 素材单查(用于拿到 file_url 传给对口型等新接口 ── */
export const getAssetById = async (id: string): Promise<{ file_url?: string; id: string }> => {
const response = await apiClient.get<{ file_url?: string; id: string }>(`/assets/${id}`)
return response.data
}
/* ── 对口型(模式A:TTS 直生,后端内部合成音频;不要先调 TTS 拿 audio_url ── */
/* ── 对口型 ── */
export const createLipsyncJob = async (data: {
/** 人物视频 URLMP4);由素材 id 经 getAssetById 拿 file_url,禁止传 video_asset_id */
video_url: string
/** 音色 ID(预置音色 或 克隆音色 profile UUID,后端会解析) */
voice_id: string
/** 要合成的文案(手动输入或文案库内容) */
script_text: string
/** 语速 0.5~2.0,默认 1.0 */
speed?: number
/** 情绪英文枚举:natural/excited/calm/friendly */
emotion?: string
enable_video_loop?: boolean
project_id?: string
video_url: string
}): Promise<LipsyncJob> => {
const response = await apiClient.post<LipsyncJob>("/lipsync/jobs", data)
return response.data
@@ -58,18 +49,6 @@ export const getLipsyncJob = async (id: string): Promise<LipsyncJob> => {
return response.data
}
/* ── 智能封面(MediaKit 抽帧 + 质量评分选最佳帧,独立于渲染任务) ── */
export const generateSmartCover = async (
video_url: string,
max_frames = 5,
): Promise<{ cover_url: string; status: string; message: string }> => {
const response = await apiClient.post<{ cover_url: string; status: string; message: string }>(
"/ai-avatar/render/smart-cover",
{ video_url, max_frames },
)
return response.data
}
/* ── 渲染 ── */
export const submitRender = async (data: {
lipsync_job_id: string
@@ -17,10 +17,6 @@ interface PanelCoverAndGenerateProps {
onResolutionChange: (r: string) => void
isGenerating: boolean
onGenerate: () => void
/** 智能获取封面(MediaKit 选帧) */
onSmartCover: () => void
smartCoverLoading: boolean
canSmartCover: boolean
/** 配置汇总信息 */
summary: {
videoName: string | null
@@ -54,9 +50,6 @@ const PanelCoverAndGenerate: React.FC<PanelCoverAndGenerateProps> = ({
onResolutionChange,
isGenerating,
onGenerate,
onSmartCover,
smartCoverLoading,
canSmartCover,
summary,
}) => {
const uploadInputRef = useRef<HTMLInputElement>(null)
@@ -76,10 +69,9 @@ const PanelCoverAndGenerate: React.FC<PanelCoverAndGenerateProps> = ({
e.target.value = ""
}
/** 智能获取封面(调后端 MediaKit 抽帧评分选最佳帧,#1822 */
const handleSmartCover = () => {
/** 从视频截取(使用配置的帧时间,默认首帧 */
const handleCaptureFromVideo = () => {
onCoverConfigChange({ mode: "auto_frame" })
onSmartCover()
}
const lipsync = summary.lipsyncStatus ? LIPSYNC_STATUS_LABEL[summary.lipsyncStatus] : null
@@ -101,11 +93,9 @@ const PanelCoverAndGenerate: React.FC<PanelCoverAndGenerateProps> = ({
<button
type="button"
className={`aa-btn aa-btn--ghost${coverConfig.mode === "auto_frame" ? " active" : ""}`}
onClick={handleSmartCover}
disabled={smartCoverLoading || !canSmartCover}
title={canSmartCover ? "基于对口型成片智能选帧" : "请先完成对口型生成"}
onClick={handleCaptureFromVideo}
>
{smartCoverLoading ? "⏳ 智能选帧中…" : "🎬 智能获取封面"}
🎬
</button>
<button
type="button"
@@ -7,17 +7,16 @@
* - AiAvatarTitleConfig ↔ TitleSettings 的双向适配
* - 自动生成字幕开关
*/
import React, { useMemo, useState, useEffect } from "react"
import { Input } from "antd"
import React, { useMemo, useState } from "react"
import TitleStylePanel from "@/pages/generate/components/title/TitleStylePanel"
import TitleLibraryAutoComplete from "@/pages/generate/components/title/TitleLibraryAutoComplete"
import type { TitleOption } from "@/pages/generate/components/title/TitleLibraryAutoComplete"
import type { TitleSettings } from "@/pages/generate/types"
import { POSITION_OPTIONS, FONT_OPTIONS, TITLE_PRESETS } from "@/pages/generate/constants"
import {
POSITION_OPTIONS,
FONT_OPTIONS,
TITLE_PRESETS,
getFontFamily,
} from "@/pages/generate/constants"
import type { AiAvatarTitleConfig } from "../types"
import { getTitles } from "@/api/titles"
const { TextArea } = Input
interface PanelTitleConfigProps {
titleConfig: AiAvatarTitleConfig
@@ -28,14 +27,6 @@ const PanelTitleConfig: React.FC<PanelTitleConfigProps> = ({ titleConfig, onUpda
/** TitleStylePanel 内部高亮的预设 key(面板本地状态) */
const [activePreset, setActivePreset] = useState<string | null>(null)
/** 标题库选项(复用智能剪辑的标题库) */
const [titleOptions, setTitleOptions] = useState<TitleOption[]>([])
useEffect(() => {
getTitles()
.then((items) => setTitleOptions(items.map((t) => ({ label: t.content, value: t.content }))))
.catch(() => setTitleOptions([]))
}, [])
/** AiAvatarTitleConfig → TitleSettings(补齐 aiAutoSelect / 自由坐标字段) */
const titleSettings: TitleSettings = useMemo(
() => ({
@@ -71,32 +62,37 @@ const PanelTitleConfig: React.FC<PanelTitleConfigProps> = ({ titleConfig, onUpda
return (
<div className="aa-title-config">
{/* 主标题输入 — TextArea 多行 + 标题库选择 */}
{/* 主标题输入 */}
<div className="aa-form-field">
<label className="aa-label"></label>
<TextArea
className="aa-title-input"
placeholder="输入视频标题(支持 / 分行)"
<input
className="aa-input aa-title-input"
type="text"
placeholder="输入视频标题(留空则不显示标题)"
value={titleConfig.title}
autoSize={{ minRows: 2, maxRows: 4 }}
maxLength={200}
maxLength={30}
onChange={(e) => onUpdate({ title: e.target.value })}
style={{ fontSize: 15 }}
/>
<div style={{ marginTop: 8, display: "flex", alignItems: "center", gap: 8 }}>
<span style={{ fontSize: 12, color: "#8c8ca1", whiteSpace: "nowrap" }}>📚 </span>
<TitleLibraryAutoComplete
key={titleConfig.title}
placeholder="选择标题填入上方"
value=""
onChange={(val) => {
if (val) onUpdate({ title: val })
{titleConfig.title && (
<div
style={{
fontSize: 13,
padding: "6px 8px",
background: "#f8f8fc",
borderRadius: 6,
fontFamily: getFontFamily(titleConfig.font),
fontWeight: titleConfig.bold ? 700 : 400,
fontStyle: titleConfig.italic ? "italic" : "normal",
color: titleConfig.color,
textShadow: titleConfig.shadow ? "1px 1px 3px rgba(0,0,0,0.6)" : undefined,
overflow: "hidden",
textOverflow: "ellipsis",
whiteSpace: "nowrap",
}}
options={titleOptions}
maxLength={200}
style={{ flex: 1 }}
/>
</div>
>
{titleConfig.title}
</div>
)}
</div>
{/* 标题样式:直接复用智能剪辑 TitleStylePanel(位置/字体/字号/样式/预设) */}
@@ -4,15 +4,12 @@
* - 已选视频:竖屏 9:16 预览播放器 + 视频信息卡片 + 移除按钮
*/
import type { AssetItem } from "@/api/assets"
import type { AiAvatarTitleConfig } from "../types"
import { getFontFamily } from "@/pages/generate/constants"
export interface PanelVideoSelectorProps {
selectedVideo: AssetItem | null
/** 触发打开素材库弹窗 */
onSelectVideo: () => void
onRemoveVideo: () => void
titleConfig?: AiAvatarTitleConfig
}
/** 格式化时长(秒 → mm:ss */
@@ -27,7 +24,6 @@ export function PanelVideoSelector({
selectedVideo,
onSelectVideo,
onRemoveVideo,
titleConfig,
}: PanelVideoSelectorProps) {
/* 未选视频:虚线上传区,点击打开素材库弹窗 */
if (!selectedVideo) {
@@ -57,42 +53,13 @@ export function PanelVideoSelector({
return (
<div>
{/* 竖屏 9:16 视频预览播放器 + 标题实时预览 */}
<div className="aa-video-preview" style={{ position: "relative" }}>
{/* 竖屏 9:16 视频预览播放器 */}
<div className="aa-video-preview">
{fileUrl ? (
<video src={fileUrl} poster={selectedVideo.thumbnail_url} controls playsInline />
) : (
<div className="aa-video-preview__placeholder"></div>
)}
{titleConfig?.title && (
<div
style={{
position: "absolute",
left: "50%",
transform: "translateX(-50%)",
...(titleConfig.position === "top"
? { top: "10%" }
: titleConfig.position === "bottom"
? { bottom: "10%" }
: { top: "50%", transform: "translate(-50%, -50%)" }),
fontSize: Math.max(titleConfig.size, 32),
fontFamily: getFontFamily(titleConfig.font),
color: titleConfig.color,
fontWeight: titleConfig.bold ? 700 : 400,
fontStyle: titleConfig.italic ? "italic" : "normal",
textShadow: "0 2px 4px rgba(0,0,0,0.5)",
WebkitTextStroke: "2px #000",
pointerEvents: "none",
zIndex: 10,
maxWidth: "90%",
textAlign: "center",
whiteSpace: "pre-wrap",
lineHeight: 1.3,
}}
>
{titleConfig.title}
</div>
)}
</div>
{/* 视频信息卡片:文件名 / 时长 / 分辨率 */}
@@ -6,7 +6,6 @@ import { useEffect, useRef, useState } from "react"
import { message } from "antd"
import { fetchVoices } from "@/api/voices/voices"
import { previewTts } from "@/api/tts"
import { normalizeEmotion } from "../utils/contract"
import type { UnifiedVoiceItem } from "@/api/voices/types"
import {
type VoiceSource,
@@ -93,14 +92,11 @@ export function PanelVoiceSelector({
/** 用指定 URL 真实播放(抽取公共) */
const playAudioUrl = (voiceId: string, url: string) => {
// 临时兼容:后端 /tts/preview 返回 HTTP URLstaging 是 HTTPSMixed Content 会阻止加载
// OSS 同时支持 HTTP/HTTPS,直接替换协议即可
const safeUrl = url.startsWith("http://") ? url.replace("http://", "https://") : url
if (audioRef.current) {
audioRef.current.pause()
audioRef.current = null
}
const audio = new Audio(safeUrl)
const audio = new Audio(url)
audioRef.current = audio
setPreviewingId(voiceId)
audio.onended = () => {
@@ -151,8 +147,7 @@ export function PanelVoiceSelector({
const res = await previewTts({
text: VOICE_PREVIEW_TEXT,
voice_id: targetId,
speed: speed, // 透传用户选择的语速(#1822)
emotion: normalizeEmotion(emotion), // 情绪中文→英文枚举
speed: 1.0,
})
console.log("[AI数字人-克隆试听] previewTts 响应:", {
audio_url: res.audio_url?.substring(0, 80),
@@ -252,7 +247,7 @@ export function PanelVoiceSelector({
type="button"
className="aa-voice-card__preview"
title={previewingId === voice.id ? "停止试听" : "试听"}
disabled={voice.type === "preset" && !previewUrl}
disabled={!previewUrl}
onClick={(e) => {
e.stopPropagation()
handlePreview(voice)
@@ -1,97 +0,0 @@
/**
* AI数字人 — 标题库选择弹窗
* 复用智能剪辑的标题库 API,选择标题后填入输入框
*/
import React, { useEffect, useState } from "react"
import { getTitles } from "@/api/titles"
import type { TitleItem } from "@/api/titles/types"
interface TitleLibraryModalProps {
open: boolean
onClose: () => void
onSelect: (title: string) => void
}
const TitleLibraryModal: React.FC<TitleLibraryModalProps> = ({ open, onClose, onSelect }) => {
const [titles, setTitles] = useState<TitleItem[]>([])
const [loading, setLoading] = useState(false)
const [search, setSearch] = useState("")
useEffect(() => {
if (!open) return
setLoading(true)
getTitles()
.then((items) => setTitles(items))
.catch(() => setTitles([]))
.finally(() => setLoading(false))
}, [open])
const filtered = titles.filter(
(t) => !search || t.content.toLowerCase().includes(search.toLowerCase()),
)
if (!open) return null
return (
<div className="aa-modal-overlay" onClick={onClose}>
<div className="aa-modal" onClick={(e) => e.stopPropagation()} style={{ maxWidth: 600 }}>
<div className="aa-modal__header">
<span className="aa-modal__title"></span>
<button className="aa-modal__close" onClick={onClose}></button>
</div>
<div className="aa-modal__body">
<div style={{ marginBottom: 12 }}>
<input
className="aa-input"
placeholder="搜索标题..."
value={search}
onChange={(e) => setSearch(e.target.value)}
/>
</div>
{loading ? (
<div style={{ textAlign: "center", padding: 40, color: "#8c8ca1" }}>...</div>
) : filtered.length === 0 ? (
<div style={{ textAlign: "center", padding: 40, color: "#8c8ca1" }}>
</div>
) : (
<div style={{ maxHeight: 400, overflowY: "auto" }}>
{filtered.map((t) => (
<div
key={t.id}
style={{
padding: "12px 16px",
marginBottom: 8,
background: "#f8f8fc",
borderRadius: 8,
cursor: "pointer",
transition: "background 0.2s",
}}
onMouseEnter={(e) => (e.currentTarget.style.background = "#eef0ff")}
onMouseLeave={(e) => (e.currentTarget.style.background = "#f8f8fc")}
onClick={() => {
onSelect(t.content)
onClose()
}}
>
<div style={{ fontSize: 14, color: "#1a1a2e", marginBottom: 4 }}>{t.content}</div>
<div style={{ fontSize: 12, color: "#8c8ca1" }}>
{t.word_count ?? t.content.length} ·{" "}
{t.created_at ? new Date(t.created_at).toLocaleDateString() : ""}
</div>
</div>
))}
</div>
)}
</div>
<div className="aa-modal__footer">
<button className="aa-btn" onClick={onClose}>
</button>
</div>
</div>
</div>
)
}
export default TitleLibraryModal
+1 -9
View File
@@ -77,9 +77,6 @@ export interface AiAvatarTitleConfig {
shadow: boolean
color: string
auto_subtitle: boolean
/** 自定义位置坐标(position=custom 时生效,像素) */
pos_x?: number
pos_y?: number
}
/* ── 封面配置 ── */
@@ -89,8 +86,6 @@ export interface AiAvatarCoverConfig {
frame_time: number
upload_url: string | null
thumbnail_url: string | null
/** 智能封面(MediaKit 选帧)返回的 OSS 非临时 URL#1822 */
smart_cover_url: string | null
}
/* ── 渲染任务 ── */
@@ -108,7 +103,7 @@ export interface RenderJob {
/* ── 默认值 ── */
export const DEFAULT_TITLE_CONFIG: AiAvatarTitleConfig = {
title: "",
position: "bottom",
position: "top",
font: "思源黑体",
size: 28,
bold: true,
@@ -117,8 +112,6 @@ export const DEFAULT_TITLE_CONFIG: AiAvatarTitleConfig = {
shadow: false,
color: "#ffffff",
auto_subtitle: true,
pos_x: undefined,
pos_y: undefined,
}
export const DEFAULT_COVER_CONFIG: AiAvatarCoverConfig = {
@@ -127,5 +120,4 @@ export const DEFAULT_COVER_CONFIG: AiAvatarCoverConfig = {
frame_time: 0,
upload_url: null,
thumbnail_url: null,
smart_cover_url: null,
}
@@ -1,76 +0,0 @@
/**
* AI数字人 — 前后端接口契约转换工具(#1822)
*
* 以 packages/domain/video_filter_builder.py 的 build_title_drawtext_filter() 为唯一口径
* (契约文档第 5 节的 titles[]/fontSize/frame/start/end 为误写,后端不认,禁止使用)。
*/
import type { AiAvatarTitleConfig, AiAvatarCoverConfig, VoiceEmotion } from "../types"
/* ── 情绪:中文 → 英文(防御性映射;state 默认已是英文) ── */
const EMOTION_ZH_TO_EN: Record<string, VoiceEmotion> = {
: "natural",
: "excited",
: "calm",
: "friendly",
}
const VALID_EMOTIONS: VoiceEmotion[] = ["natural", "excited", "calm", "friendly"]
/** 归一化为后端英文枚举 natural/excited/calm/friendly;非法/空值回退 natural。 */
export function normalizeEmotion(raw: string | undefined | null): VoiceEmotion {
if (!raw) return "natural"
const v = raw.trim()
if ((VALID_EMOTIONS as string[]).includes(v)) return v as VoiceEmotion
return EMOTION_ZH_TO_EN[v] ?? "natural"
}
/* ── 标题:前端 state → 后端 build_title_drawtext_filter 字段(单个 title_config dict ── */
/**
* 后端真实字段:text(或content)、font(或font_preset)、font_size(或size)、
* font_color(或color,可传 #RRGGBB)、position(top/center/bottom/custom)、
* enabled、bold、stroke{enabled,width,color}、shadow{enabled,color,offset_x,offset_y}、
* pos_x/pos_y(custom 时)。
* 口播标题默认 position=bottom(不传后端会默认 top 跑到画面顶部)。
*/
export function buildTitleConfigPayload(cfg: AiAvatarTitleConfig): Record<string, unknown> {
const text = (cfg.title || "").trim()
if (!text) return {}
const position = cfg.position || "bottom"
const payload: Record<string, unknown> = {
text,
enabled: true,
font: cfg.font || "思源黑体",
font_size: Math.round(cfg.size) || 36,
font_color: cfg.color || "#ffffff",
position,
bold: !!cfg.bold,
stroke: cfg.stroke ? { enabled: true, width: 2, color: "#000000" } : { enabled: false },
shadow: cfg.shadow
? { enabled: true, color: "#000000", offset_x: 2, offset_y: 2 }
: { enabled: false },
}
// 自定义坐标(custom 位置)
if (position === "custom" && typeof cfg.pos_x === "number" && typeof cfg.pos_y === "number") {
payload.pos_x = cfg.pos_x
payload.pos_y = cfg.pos_y
}
return payload
}
/* ── 封面:前端 state → 后端 render cover_config ── */
export function buildCoverConfigPayload(
cfg: AiAvatarCoverConfig,
smartCoverUrl: string | null,
): Record<string, unknown> {
const payload: Record<string, unknown> = {
enabled: !!cfg.enabled,
mode: cfg.mode,
// build_cover_extract_command 读取 timestamp(截帧秒数)
timestamp: cfg.frame_time || 0,
}
if (smartCoverUrl) payload.cover_url = smartCoverUrl
// 自定义上传:blob: 本地预览地址无法给后端,仅 OSS URL 可用
if (cfg.mode === "upload" && cfg.upload_url && !cfg.upload_url.startsWith("blob:")) {
payload.upload_url = cfg.upload_url
}
return payload
}
-169
View File
@@ -1,169 +0,0 @@
# AI 数字人前后端接口契约(#1797 / #1822
> 分支:`fix/ai-avatar-v3-1797`
> 范围:TTS→对口型链路打通、语速/情绪透传、封面智能选帧、标题字段对齐
> 本文档为前后端联调的唯一字段口径。
---
## 1. 对口型创建接口 `POST /api/v1/lipsync/jobs`
支持两种输入模式,**二选一**
### 模式 A(推荐):TTS 直生 —— 传音色 + 文案,后端内部合成音频
前端无需先调 TTS。后端收到请求后:先调 CosyVoice 合成音频 → 转存 OSS → 再提交 MediaKit 对口型。
```jsonc
{
"video_url": "https://oss.../person.mp4", // 必填,人物视频(MP4
"voice_id": "cosyvoice-v3-flash-99-xxxx", // 必填,音色 ID(预置音色 或 克隆 profile UUID
"script_text": "省是浙江省,市是永康市……", // 必填,要合成的文案
"speed": 1.0, // 可选,语速 0.5~2.0,默认 1.0
"emotion": "excited", // 可选,情绪,见 §3
"enable_video_loop": false, // 可选,音频长于视频时是否循环画面
"project_id": "" // 可选
}
```
### 模式 B:直接音频 —— 前端已准备好音频
```jsonc
{
"video_url": "https://oss.../person.mp4", // 必填
"audio_url": "https://oss.../voice.mp3", // 必填,mp3/aac/wav/m4a/flac
"enable_video_loop": false
}
```
### 校验与错误码
| 场景 | HTTP | detail.code |
|------|------|-------------|
| 既无 audio_url 又无 voice_id+script_text | 422 | schema 校验) |
| video_url 非 MP4 / audio_url 格式不支持 | 422 | schema 校验) |
| 克隆音色不属于当前用户 | 403 | `VoiceForbidden` |
| 克隆音色尚未合成完成 | 400 | `VoiceNotReady` |
| TTS 合成失败(如 CosyVoice 欠费) | 502 | `TTSSynthesisFailed` |
| MediaKit 提交失败 | 502 | `*`(透传 MediaKit code |
### 轮询
- `GET /api/v1/lipsync/jobs/{id}`:非终态任务先返回 DB 缓存,**后台异步刷新 MediaKit**(不会阻塞轮询)。
- `status` 流转:`pending``submitted``running`/`processing`MediaKit 中间态同步)→ `completed` / `failed`
- `completed``output_video_url` 为**已转存自家 OSS 的非临时 URL**(不会过期)。
- 前端每 3s 轮询,命中 `completed`/`failed` 即停。
---
## 2. TTS 合成接口语速/情绪透传
- `POST /api/v1/tts/synthesize`(异步任务)与 `POST /api/v1/tts/preview`(即时试听)均新增:
- `speed`float0.5~2.0,默认 1.0 → 透传 CosyVoice payload 的 `rate`
- `emotion`:string,见 §3 映射 → 透传 `emotion`
- 透传链路:`route → CreateTTSJobUseCase(metadata) → TTSJobWorkflow.start_synthesis / 分段合成 → CosyVoiceService.submit_synthesize_task(rate/emotion)`
- 分段合成(长文案)与失败重合成路径同样透传 speed/emotion。
---
## 3. 情绪枚举(前后端统一)
前端把中文选项映射成英文后传后端;后端同时接受中文/英文,非法值忽略(走默认自然)。
| 前端选项 | 传参值 | CosyVoice 枚举 |
|---------|--------|---------------|
| 自然 | `natural` | natural |
| 兴奋 | `excited` | excited |
| 沉稳 | `calm` | calm |
| 亲切 | `friendly` | friendly |
后端 `normalize_emotion()` 也接受中文(自然/兴奋/沉稳/亲切)做兜底映射。
---
## 4. 智能封面接口 `POST /api/v1/ai-avatar/render/smart-cover`
独立接口,**不依赖渲染任务**,前端「智能获取封面」按钮直接调用。
**请求**
```jsonc
{
"video_url": "https://oss.../avatar_output.mp4", // 必填,数字人视频
"max_frames": 5 // 可选,抽帧数量 1~10,默认 5
}
```
**响应**
```jsonc
{
"cover_url": "https://oss.../ai-avatar/covers/xxx/cover_yy.jpg", // OSS 非临时 URL
"status": "completed", // completed / fallback_failed
"message": "" // 失败原因
}
```
**实现**:复用智能剪辑同款能力 —— MediaKit `extract_frames(SpecifiedFrames)` 抽 5 帧 → `cover_frame_scorer.score_frames`(清晰度+亮度+色彩)评分选最佳 → 转存 OSS。
**不再使用 FFmpeg 简单首帧**。渲染管线最终封面也优先走该智能选帧,MediaKit 不可用时才回退 FFmpeg。
---
## 5. 渲染接口 `POST /api/v1/ai-avatar/render`
```jsonc
{
"lipsync_job_id": "7c29a3b2-...", // 必填,已 completed 的对口型任务
"script_id": "", // 可选!见下方说明
"b_roll_segments": [], // 可选,B-roll 片段
"title_config": { ... }, // 可选,单个标题配置 dict(见 §6)
"cover_config": { ... }, // 可选,封面配置(建议改用 smart-cover
"project_id": ""
}
```
**`script_id` 是否必填:可选。**
- 从文案库选了文案时传对应文案 ID(后端做归属校验)。
- **手动输入文案、走 TTS 直生模式时不传(留空)即可**——渲染管线不依赖文案内容,`script_id` 仅用于归属校验。留空不会卡手动文案用户。
---
## 6. 标题配置 `title_config` 字段清单(以 build_title_drawtext_filter 为准)
渲染请求收的是**单个 `title_config` dict**(不是 `titles[]` 数组),字段与 `packages/domain/video_filter_builder.py``build_title_drawtext_filter()` 完全对齐:
| 字段 | 别名 | 类型 | 必填 | 默认 | 说明 |
|------|------|------|------|------|------|
| `text` | `content` | string | ✅ | — | 标题文字;为空或 `enabled=false` 时不渲染标题 |
| `enabled` | — | bool | ❌ | `true` | 是否启用标题;false 跳过 |
| `font` | `font_preset` | string | ❌ | 思源黑体 | 字体名(后端按名字解析字体文件) |
| `font_size` | `size` | int | ❌ | 36 | 字号(像素) |
| `font_color` | `color` | string | ❌ | `#ffffff` | 文字颜色,`#RRGGBB`;后端自动去掉 `#`,也可传 `RRGGBB` 或颜色名 |
| `position` | — | string | ❌ | `top` | 预设位置:`top`(y=50) / `center`(垂直居中) / `bottom`(底部上移50px) / `custom` |
| `pos_x` | — | int/float | ❌ | — | 自定义 X 坐标(像素),仅 `position=custom` 生效 |
| `pos_y` | — | int/float | ❌ | — | 自定义 Y 坐标(像素),仅 `position=custom` 生效 |
| `bold` | — | bool | ❌ | `true` | 粗体(Bold 字体变体,回退 borderw 模拟) |
| `stroke` | — | bool/object | ❌ | — | 描边。`true`=黑描边宽2object 见下 |
| `stroke.enabled` | — | bool | ❌ | true | 是否描边 |
| `stroke.width` | — | int | ❌ | 2 | 描边宽度 |
| `stroke.color` | — | string | ❌ | `#000000` | 描边颜色 |
| `shadow` | — | bool/object | ❌ | — | 阴影。`true`=黑色阴影偏移2pxobject 见下 |
| `shadow.enabled` | — | bool | ❌ | true | 是否阴影 |
| `shadow.color` | — | string | ❌ | `#000000` | 阴影颜色 |
| `shadow.offset_x` | — | int | ❌ | 2 | 阴影 X 偏移 |
| `shadow.offset_y` | — | int | ❌ | 2 | 阴影 Y 偏移 |
**前端注意事项**
- 标题是**整条成片一个标题**(单个 dict),不是按时间段的标题数组;没有 `start`/`end`/`frame`/`fontSize` 这些字段。
- 位置用 `position` 四档枚举;自由摆放用 `position="custom"` + `pos_x`/`pos_y`(像素坐标,非比例)。
- 颜色统一传 `#RRGGBB` 即可,后端会处理 `#`;三档预设位置下标题始终水平居中。
- `stroke`/`shadow``true` 用默认样式,或传 object 精细控制颜色/宽度/偏移。
---
## 7. 前端对接清单
1. 对口型:改用**模式 A**voice_id + script_text + speed + emotion),不要再先调 TTS 拿 audio_url。
2. 音色 ID`voice_id` 可直接传克隆音色的 profile UUID,后端会解析为 CosyVoice voice_id(与 /tts 一致)。
3. 情绪下拉:自然/兴奋/沉稳/亲切 → natural/excited/calm/friendly。
4. 封面:点「智能获取封面」→ POST `/ai-avatar/render/smart-cover`,用返回的 `cover_url`
5. 渲染:手动文案直生场景 `script_id` 留空;标题传**单个** `title_config` dict(字段见 §6)。
6. 轮询:识别 `running` 等中间态,不要只认 `submitted`
+2 -2
View File
@@ -128,9 +128,9 @@ services:
- xiaoxia-net
# 健康检查配置
# 注:容器内无 pgrep/ps,扫描 /proc 所有进程的 cmdline 查找 celery 进程
# 注:celery inspect ping 依赖 broker 连接,在容器内不可靠,改用进程检查
healthcheck:
test: ["CMD-SHELL", "grep -lq celery /proc/[0-9]*/cmdline 2>/dev/null || exit 1"]
test: ["CMD-SHELL", "for pid in /proc/[0-9]*/cmdline; do if grep -ql celery \"$pid\" 2>/dev/null; then exit 0; fi; done; exit 1"]
interval: 30s
timeout: 10s
retries: 3
+1 -1
View File
@@ -155,7 +155,7 @@ docker run -d \
--restart unless-stopped \
--cpus 2 \
--memory 2g \
--health-cmd "sh -c \"grep -lq celery /proc/[0-9]*/cmdline 2>/dev/null || exit 1\"" \
--health-cmd "sh -c \"for pid in /proc/[0-9]*/cmdline; do if grep -ql celery \"$pid\" 2>/dev/null; then exit 0; fi; done; exit 1\"" \
--health-interval 30s \
--health-timeout 10s \
--health-retries 3 \
+1 -1
View File
@@ -116,7 +116,7 @@ docker run -d \
-v "$GENERATED_DIR:/app/generated" \
--restart unless-stopped \
--label com.centurylinklabs.watchtower.enable=true \
--health-cmd "sh -c \"grep -lq celery /proc/[0-9]*/cmdline 2>/dev/null || exit 1\"" \
--health-cmd "sh -c \"for pid in /proc/[0-9]*/cmdline; do if grep -ql celery \"$pid\" 2>/dev/null; then exit 0; fi; done; exit 1\"" \
--health-interval 30s \
--health-timeout 10s \
--health-retries 3 \
-4
View File
@@ -38,10 +38,6 @@ RUN chmod +x /usr/local/bin/entrypoint-worker.sh
# 业务代码(变化最频繁,放最后)
COPY apps/worker/ /app/apps/worker/
# 健康检查:扫描所有进程的 cmdline 查找 celery 进程
HEALTHCHECK --interval=30s --timeout=10s --start-period=40s --retries=3 \
CMD grep -lq celery /proc/[0-9]*/cmdline 2>/dev/null || exit 1
USER celery
WORKDIR /app/apps/worker
CMD ["/usr/local/bin/entrypoint-worker.sh"]
+2 -9
View File
@@ -684,15 +684,9 @@ class LipsyncJobModel(Base):
# 输入参数
video_url = Column(Text, nullable=False)
audio_url = Column(Text, nullable=True) # 直生模式(voice_id+script_text)下 TTS 合成后回填
audio_url = Column(Text, nullable=False)
enable_video_loop = Column(Boolean, nullable=False, default=False)
# TTS 直生字段:传音色 + 文案,由后端先合成音频再对口型
voice_id = Column(String(200), nullable=False, default="")
script_text = Column(Text, nullable=False, default="")
speed = Column(Float, nullable=False, default=1.0)
emotion = Column(String(20), nullable=False, default="")
# MediaKit 任务状态
mediakit_task_id = Column(String(200), nullable=False, default="", index=True)
status = Column(
@@ -721,8 +715,7 @@ class AiAvatarRenderJob(Base):
# 输入参数
lipsync_job_id = Column(String(36), nullable=False)
# 文案 ID 可选:手动输入文案(TTS 直生)场景不关联文案库条目
script_id = Column(String(36), nullable=False, default="")
script_id = Column(String(36), nullable=False)
b_roll_segments = Column(JSON, nullable=False, default=list)
# b_roll_segments 格式: [{"script_segment_index": 0, "asset_url": "...", "mode": "fullscreen|pip", "start_time": 5.0, "end_time": 10.0}, ...]
title_config = Column(JSON, nullable=False, default=dict)
+10 -50
View File
@@ -25,35 +25,6 @@ from packages.shared.config import get_shared_settings
logger = logging.getLogger(__name__)
# CosyVoice 支持的情绪:中文标签 → API 英文值
EMOTION_MAP = {
"自然": "natural",
"兴奋": "excited",
"沉稳": "calm",
"亲切": "friendly",
"natural": "natural",
"excited": "excited",
"calm": "calm",
"friendly": "friendly",
}
VALID_EMOTIONS = {"natural", "excited", "calm", "friendly"}
def normalize_emotion(emotion: str) -> str:
"""将前端情绪值归一化为 CosyVoice 英文枚举。
支持中文(自然/兴奋/沉稳/亲切)和英文;非法值返回空串(不传,走默认)。
"""
if not emotion:
return ""
key = emotion.strip().lower()
mapped = EMOTION_MAP.get(emotion.strip()) or EMOTION_MAP.get(key)
if mapped and mapped in VALID_EMOTIONS:
return mapped
logger.warning("未知的 emotion 值,忽略: %r", emotion)
return ""
class CosyVoiceError(Exception):
"""CosyVoice API 调用异常。"""
@@ -460,7 +431,6 @@ class CosyVoiceService:
format: str = "",
speed: float = 1.0,
volume: int = 50,
emotion: str = "",
) -> dict:
"""提交语音合成任务(同步非流式,直接返回结果).
@@ -474,7 +444,6 @@ class CosyVoiceService:
format: 输出格式(mp3/wav/pcm),空表示使用配置默认值
speed: 语速(0.5-2.0),1.0 为正常速度
volume: 音量(0-100),默认 50
emotion: 情绪(natural/excited/calm/friendly),空串不传
Returns:
dict: {"audio_url": str, "request_id": str,
@@ -494,20 +463,17 @@ class CosyVoiceService:
settings = get_shared_settings()
input_payload: dict[str, Any] = {
"text": text,
"voice": voice_id,
"format": format or settings.cosyvoice_format,
"sample_rate": sample_rate or settings.cosyvoice_sample_rate,
"rate": speed,
"volume": volume,
payload = {
"model": self._model,
"input": {
"text": text,
"voice": voice_id,
"format": format or settings.cosyvoice_format,
"sample_rate": sample_rate or settings.cosyvoice_sample_rate,
"rate": speed,
"volume": volume,
},
}
# 情绪:归一化(中文→英文)后透传;空/非法则不传,走 CosyVoice 默认
norm_emotion = normalize_emotion(emotion)
if norm_emotion:
input_payload["emotion"] = norm_emotion
payload = {"model": self._model, "input": input_payload}
response = self._call_api(
method="POST",
@@ -524,10 +490,6 @@ class CosyVoiceService:
if not audio_url:
raise CosyVoiceError(f"CosyVoice API 未返回 audio_url: {response}")
# DashScope 返回 http://,统一升级为 https://
if audio_url.startswith("http://"):
audio_url = audio_url.replace("http://", "https://", 1)
return {
"task_id": "", # 同步接口无 task_id,兼容旧接口
"audio_url": audio_url,
@@ -555,7 +517,6 @@ class CosyVoiceService:
format: str = "",
speed: float = 1.0,
volume: int = 50,
emotion: str = "",
timeout: float = 120.0,
) -> SynthesizeResult:
"""语音合成(同步非流式).
@@ -587,7 +548,6 @@ class CosyVoiceService:
format=format,
speed=speed,
volume=volume,
emotion=emotion,
)
return SynthesizeResult(
-14
View File
@@ -143,16 +143,11 @@ class TTSWorkflowService:
return self._start_segment_synthesis(job)
try:
_meta = dict(job.metadata)
_speed = float(_meta.get("speed", 1.0) or 1.0)
_emotion = str(_meta.get("emotion", "") or "")
submit_result = self.cosyvoice_service.submit_synthesize_task(
text=job.input_text,
voice_id=job.voice_id,
sample_rate=job.sample_rate,
format=job.format,
speed=_speed,
emotion=_emotion,
)
# 保存 task_id / request_id 到 metadata
@@ -288,7 +283,6 @@ class TTSWorkflowService:
job_metadata = job.metadata or {}
speed = float(job_metadata.get("speed", 1.0))
volume = int(job_metadata.get("volume", 50))
emotion = str(job_metadata.get("emotion", "") or "")
result = self.cosyvoice_service.submit_synthesize_task(
text=job.input_text,
@@ -297,7 +291,6 @@ class TTSWorkflowService:
format=job.format,
speed=speed,
volume=volume,
emotion=emotion,
)
audio_url = result.get("audio_url", "")
if not audio_url:
@@ -408,9 +401,6 @@ class TTSWorkflowService:
"""
max_workers = min(len(segments), _MAX_SEGMENT_WORKERS)
results: list[dict | None] = [None] * len(segments)
_seg_meta = job.metadata or {}
_seg_speed = float(_seg_meta.get("speed", 1.0) or 1.0)
_seg_emotion = str(_seg_meta.get("emotion", "") or "")
with ThreadPoolExecutor(max_workers=max_workers) as executor:
future_to_idx = {}
@@ -421,8 +411,6 @@ class TTSWorkflowService:
voice_id=job.voice_id,
sample_rate=job.sample_rate,
format=job.format,
speed=_seg_speed,
emotion=_seg_emotion,
)
future_to_idx[future] = idx
@@ -512,7 +500,6 @@ class TTSWorkflowService:
job_metadata = job.metadata or {}
speed = float(job_metadata.get("speed", 1.0))
volume = int(job_metadata.get("volume", 50))
emotion = str(job_metadata.get("emotion", "") or "")
# 分段文本(用于缺失段重新合成)
segments = split_text(job.input_text, max_chars=_SEGMENT_THRESHOLD)
@@ -547,7 +534,6 @@ class TTSWorkflowService:
format=job.format,
speed=speed,
volume=volume,
emotion=emotion,
)
future_to_idx[future] = idx
@@ -1,297 +0,0 @@
"""#1822 情绪/语速透传 + 对口型 TTS 直生 + 智能封面 单元测试.
CI 增量映射:
cosyvoice_service.normalize_emotion / payload emotion
lipsync_service TTS 直生分支(voice_id+script_text
ai_avatar_cover_service 智能选帧
"""
import os
from unittest.mock import MagicMock, patch
import pytest
os.environ.setdefault("JWT_SECRET_KEY", "dev-secret-key-for-testing")
# ── 情绪归一化 ──────────────────────────────────────────────────────────
def test_normalize_emotion_english_values():
from packages.application.cosyvoice_service import normalize_emotion
assert normalize_emotion("natural") == "natural"
assert normalize_emotion("excited") == "excited"
assert normalize_emotion("calm") == "calm"
assert normalize_emotion("friendly") == "friendly"
# 大小写 / 空白容错
assert normalize_emotion(" Excited ") == "excited"
def test_normalize_emotion_chinese_values():
from packages.application.cosyvoice_service import normalize_emotion
assert normalize_emotion("自然") == "natural"
assert normalize_emotion("兴奋") == "excited"
assert normalize_emotion("沉稳") == "calm"
assert normalize_emotion("亲切") == "friendly"
def test_normalize_emotion_invalid_returns_empty():
from packages.application.cosyvoice_service import normalize_emotion
assert normalize_emotion("") == ""
assert normalize_emotion("angry") == ""
assert normalize_emotion("喜怒哀乐") == ""
# ── CosyVoice payload 携带 emotion + rate ──────────────────────────────
def _make_service_with_captured_client(captured: dict):
"""构造 CosyVoiceService,拦截 post 请求体到 captured['json']."""
import httpx as _httpx
from packages.application import cosyvoice_service as mod
mock_client = MagicMock(spec=_httpx.Client)
resp = MagicMock()
resp.status_code = 200
resp.json.return_value = {
"request_id": "req-1",
"output": {"audio": {"url": "https://tts/a.mp3", "duration": 1.0}},
}
resp.raise_for_status = MagicMock()
def fake_request(method, url, headers, json, timeout):
captured["json"] = json
return resp
mock_client.request.side_effect = fake_request
with patch.object(mod, "get_shared_settings") as settings_patch:
s = MagicMock()
s.cosyvoice_api_key = "sk-test"
s.cosyvoice_base_url = "https://x/api/v1"
s.cosyvoice_model = "cosyvoice-v3-flash"
s.cosyvoice_clone_model = "voice-enrollment"
s.cosyvoice_format = "mp3"
s.cosyvoice_sample_rate = 22050
s.cosyvoice_voice = "longxiaochun"
settings_patch.return_value = s
svc = mod.CosyVoiceService(http_client=mock_client)
return svc
def test_submit_synthesize_payload_includes_emotion_and_rate():
captured: dict = {}
svc = _make_service_with_captured_client(captured)
svc.submit_synthesize_task(text="你好", voice_id="v-1", speed=1.5, emotion="兴奋")
inp = captured["json"]["input"]
assert inp["emotion"] == "excited"
assert inp["rate"] == 1.5
def test_submit_synthesize_payload_omits_emotion_when_empty():
captured: dict = {}
svc = _make_service_with_captured_client(captured)
svc.submit_synthesize_task(text="你好", voice_id="v-1")
assert "emotion" not in captured["json"]["input"]
# ── 对口型 TTS 直生分支 ─────────────────────────────────────────────────
def _lipsync_service_with_mocks():
from app.services.lipsync_service import LipsyncService
db = MagicMock()
client = MagicMock()
client.is_available = True
client.submit_lipsync.return_value = {
"success": True,
"task_id": "mk-1",
"request_id": "req-1",
}
cosy = MagicMock()
cosy.submit_synthesize_task.return_value = {
"audio_url": "https://tts/raw.mp3",
"request_id": "tts-req",
"audio_duration": 3.0,
}
svc = LipsyncService(db, client=client, cosyvoice_service=cosy, voice_clone_repo=MagicMock())
# _resolve_voice_id 默认原样返回(repo.get 返回 None
svc._voice_clone_repo.get.return_value = None
return svc, client, cosy
def test_create_job_tts_direct_mode_synthesizes_audio():
svc, client, cosy = _lipsync_service_with_mocks()
with (
patch("app.services.lipsync_service.get_shared_storage_service") as storage_patch,
patch("app.services.lipsync_service.safe_download_bytes") as dl_patch,
):
storage = MagicMock()
storage.upload_file.return_value = "https://oss/tts.mp3"
storage_patch.return_value = storage
dl_patch.return_value = b"FAKEAUDIO"
job = svc.create_job(
user_id="user-1",
video_url="https://oss/person.mp4",
voice_id="cosy-v1",
script_text="你好世界",
speed=1.2,
emotion="兴奋",
)
# 调了 TTS 合成,带 speed/emotion
cosy.submit_synthesize_task.assert_called_once()
_, kwargs = cosy.submit_synthesize_task.call_args
assert kwargs["speed"] == 1.2
assert kwargs["emotion"] == "excited"
assert kwargs["voice_id"] == "cosy-v1"
# MediaKit 用合成后的 OSS 音频 URL 提交
_, submit_kwargs = client.submit_lipsync.call_args
assert submit_kwargs["audio_url"] == "https://oss/tts.mp3"
assert submit_kwargs["video_url"] == "https://oss/person.mp4"
# DB 记录了 TTS 字段
assert job.emotion == "excited"
assert job.speed == 1.2
def test_create_job_direct_audio_mode_skips_tts():
svc, client, cosy = _lipsync_service_with_mocks()
job = svc.create_job(
user_id="user-1",
video_url="https://oss/person.mp4",
audio_url="https://oss/ready.mp3",
)
cosy.submit_synthesize_task.assert_not_called()
_, submit_kwargs = client.submit_lipsync.call_args
assert submit_kwargs["audio_url"] == "https://oss/ready.mp3"
def test_create_job_tts_failure_raises():
from app.services.mediakit_client import MediaKitError
from packages.application.cosyvoice_service import CosyVoiceError
svc, client, cosy = _lipsync_service_with_mocks()
cosy.submit_synthesize_task.side_effect = CosyVoiceError("Arrearage")
with pytest.raises(MediaKitError) as exc:
svc.create_job(
user_id="user-1",
video_url="https://oss/person.mp4",
voice_id="v-1",
script_text="文本",
)
assert exc.value.code == "TTSSynthesisFailed"
# TTS 失败不应提交 MediaKit
client.submit_lipsync.assert_not_called()
# ── refresh 同步中间状态 ────────────────────────────────────────────────
def test_refresh_syncs_running_status():
from app.services.lipsync_service import LipsyncService
db = MagicMock()
client = MagicMock()
client.get_task_status.return_value = {"success": True, "status": "running"}
svc = LipsyncService(db, client=client)
job = MagicMock()
job.status = "submitted"
job.mediakit_task_id = "mk-1"
job.id = "j-1"
svc.get_job = MagicMock(return_value=job)
result = svc.refresh_job_status("j-1", "user-1")
assert result.status == "running"
# ── 智能封面 ────────────────────────────────────────────────────────────
def test_smart_cover_selects_best_frame_and_persists():
from app.services import ai_avatar_cover_service as cov
snapshots = [
{"image_url": "https://mk/f0.jpg"},
{"image_url": "https://mk/f1.jpg"},
]
with (
patch("packages.shared.mediakit_client.get_mediakit_client") as mk_patch,
patch("packages.shared.cover_frame_scorer.score_frames") as score_patch,
patch("httpx.get") as http_get,
patch("packages.shared.storage.get_shared_storage_service") as storage_patch,
):
mk = MagicMock()
mk.is_available = True
mk.extract_frames.return_value = snapshots
mk_patch.return_value = mk
# score_frames 把 f1 选为最佳
score_patch.side_effect = lambda cands: [
{"url": "https://mk/f1.jpg", "score": 90.0, "image_path": cands[1]["image_path"]},
{"url": "https://mk/f0.jpg", "score": 60.0, "image_path": cands[0]["image_path"]},
]
resp = MagicMock()
resp.content = b"IMGDATA"
resp.raise_for_status = MagicMock()
http_get.return_value = resp
storage = MagicMock()
storage.upload_file.return_value = "https://oss/cover.jpg"
storage_patch.return_value = storage
url = cov.generate_smart_cover("https://oss/avatar.mp4", job_id="job-1")
assert url == "https://oss/cover.jpg"
mk.extract_frames.assert_called_once()
score_patch.assert_called_once()
def test_smart_cover_returns_empty_when_mediakit_unavailable():
from app.services import ai_avatar_cover_service as cov
with patch("packages.shared.mediakit_client.get_mediakit_client") as mk_patch:
mk = MagicMock()
mk.is_available = False
mk_patch.return_value = mk
url = cov.generate_smart_cover("https://oss/avatar.mp4")
assert url == ""
# ── 渲染 script_id 可选(手动文案直生场景)──────────────────────────────
def test_render_request_script_id_optional():
import os
import sys
sys.path.insert(0, os.path.join(os.path.dirname(__file__), "..", "..", "apps", "api"))
from app.schemas.ai_avatar_render import CreateAiAvatarRenderRequest
# 手动文案直生:不传 script_id 也合法
req = CreateAiAvatarRenderRequest(lipsync_job_id="j-1")
assert req.script_id == ""
# 空白被 strip
req2 = CreateAiAvatarRenderRequest(lipsync_job_id="j-1", script_id=" ")
assert req2.script_id == ""
# title_config 是单个 dict
req3 = CreateAiAvatarRenderRequest(
lipsync_job_id="j-1",
title_config={"text": "标题", "position": "top", "font_size": 40},
)
assert req3.title_config["position"] == "top"
+4 -14
View File
@@ -176,23 +176,13 @@ class TestSchemaValidation:
script_id="script-1",
)
def test_create_request_empty_script_id_normalized(self):
"""script_id 改为可选(TTS 直生场景):空白值应规范化为空串而非抛错。"""
def test_create_request_empty_script_id(self):
from app.schemas.ai_avatar_render import CreateAiAvatarRenderRequest
req = CreateAiAvatarRenderRequest(
lipsync_job_id="lipsync-1",
script_id=" ",
)
assert req.script_id == ""
def test_create_request_empty_lipsync_job_id_raises(self):
from app.schemas.ai_avatar_render import CreateAiAvatarRenderRequest
with pytest.raises(ValueError, match="lipsync_job_id 不能为空"):
with pytest.raises(ValueError, match="script_id 不能为空"):
CreateAiAvatarRenderRequest(
lipsync_job_id=" ",
script_id="script-1",
lipsync_job_id="lipsync-1",
script_id=" ",
)
+84 -206
View File
@@ -35,14 +35,14 @@ def mock_mediakit():
@pytest.fixture
def mock_cosyvoice():
"""Mock CosyVoice 服务v3: service 内部走 submit_synthesize_task,返回 dict."""
"""Mock CosyVoice 服务."""
service = MagicMock()
service.submit_synthesize_task.return_value = {
"audio_url": "https://oss.example.com/tts-output.mp3",
"request_id": "tts-req-789",
}
# synthesize_speech 保留给直接同步调用场景
service.synthesize_speech.return_value = MagicMock(audio_url="https://oss.example.com/tts-output.mp3")
service.synthesize_speech.return_value = MagicMock(
audio_url="https://oss.example.com/tts-output.mp3",
duration=15.0,
file_size=12345,
request_id="tts-req-789",
)
return service
@@ -172,121 +172,86 @@ class TestSchemaValidation:
)
assert "?token=" in req.video_url
def test_dual_mode_fields_present(self):
"""v3 契约: 双模式——支持直接音频 audio_url,也支持 TTS 直生 voice_id+script_text."""
from app.schemas.lipsync import CreateLipsyncJobRequest
fields = CreateLipsyncJobRequest.model_fields.keys()
# 直接音频模式
assert "audio_url" in fields
# TTS 直生模式
assert "voice_id" in fields
assert "script_text" in fields
# 语速/情绪透传
assert "speed" in fields
assert "emotion" in fields
def test_direct_audio_mode_accepted(self):
"""v3 契约: 只传 audio_url(直接音频模式)也合法,无需 voice_id/script_text."""
def test_no_audio_url_in_request(self):
"""#1809: 请求体不应包含 audio_url 字段."""
from app.schemas.lipsync import CreateLipsyncJobRequest
req = CreateLipsyncJobRequest(
video_url="https://example.com/video.mp4",
audio_url="https://example.com/audio.mp3",
voice_id="longxiaochun_v3",
script_text="测试文本",
)
assert req.audio_url == "https://example.com/audio.mp3"
def test_neither_mode_rejected(self):
"""v3 契约: audio_url 与 voice_id+script_text 都缺时应报错."""
from app.schemas.lipsync import CreateLipsyncJobRequest
with pytest.raises(ValueError):
CreateLipsyncJobRequest(video_url="https://example.com/video.mp4")
assert not hasattr(req, "audio_url")
fields = req.model_fields.keys()
assert "audio_url" not in fields
assert "voice_id" in fields
assert "script_text" in fields
class TestLipsyncServiceUnit:
"""Service 层单元测试(纯 mock,不依赖数据库)— #1809 更新."""
def test_create_job_success(self, mock_mediakit, mock_cosyvoice):
"""v3: TTS 直生——service 内部 submit_synthesize_task 合成后转存 OSS,再提交 MediaKit."""
from app.services.lipsync_service import LipsyncService
mock_db = MagicMock()
mock_repo = MagicMock()
mock_repo.get.return_value = None # 预置音色,原样返回 voice_id
mock_db.add = MagicMock()
mock_db.flush = MagicMock()
mock_db.commit = MagicMock()
mock_db.refresh = MagicMock()
with (
patch("app.services.lipsync_service.get_shared_storage_service") as storage_patch,
patch("app.services.lipsync_service.safe_download_bytes") as dl_patch,
):
storage_patch.return_value.upload_file.return_value = "https://my-oss/tts.mp3"
dl_patch.return_value = b"audio-bytes"
svc = LipsyncService(mock_db, client=mock_mediakit, cosyvoice_service=mock_cosyvoice)
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=mock_repo,
)
job = svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="longxiaochun_v3",
script_text="大家好,欢迎来到直播间",
speed=1.2,
emotion="兴奋",
)
job = svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="longxiaochun_v3",
script_text="大家好,欢迎来到直播间",
)
assert job.status == "submitted"
assert job.mediakit_task_id == "mk-task-123"
# TTS 直生走 submit_synthesize_task,带语速/情绪
mock_cosyvoice.submit_synthesize_task.assert_called_once()
_, kwargs = mock_cosyvoice.submit_synthesize_task.call_args
assert kwargs["text"] == "大家好,欢迎来到直播间"
assert kwargs["voice_id"] == "longxiaochun_v3"
assert kwargs["speed"] == 1.2
assert kwargs["emotion"] == "excited" # 兴奋→excited
# job 记录透传字段
assert job.speed == 1.2
assert job.emotion == "excited"
# MediaKit 用转存后的 OSS audio_url
# TTS 应该被调用
mock_cosyvoice.synthesize_speech.assert_called_once_with(
text="大家好,欢迎来到直播间",
voice_id="longxiaochun_v3",
)
# MediaKit 应该用 TTS 生成的 audio_url
mock_mediakit.submit_lipsync.assert_called_once()
call_kwargs = mock_mediakit.submit_lipsync.call_args
assert call_kwargs.kwargs["audio_url"] == "https://my-oss/tts.mp3"
assert call_kwargs.kwargs["audio_url"] == "https://oss.example.com/tts-output.mp3"
def test_create_job_tts_failure(self, mock_mediakit):
"""v3: TTS 合成失败时,CosyVoiceError 被包装为 MediaKitError(TTSSynthesisFailed)
在建 DB 记录之前抛出,不提交 MediaKit。"""
"""TTS 合成失败时,应创建 failed 记录并抛出 CosyVoiceError."""
from app.services.lipsync_service import LipsyncService
from app.services.mediakit_client import MediaKitError
from packages.application.cosyvoice_service import CosyVoiceError
mock_cosyvoice = MagicMock()
mock_cosyvoice.submit_synthesize_task.side_effect = CosyVoiceError("Arrearage 欠费")
mock_cosyvoice.synthesize_speech.side_effect = CosyVoiceError("TTS 服务不可用")
mock_db = MagicMock()
mock_repo = MagicMock()
mock_repo.get.return_value = None
mock_db.add = MagicMock()
mock_db.flush = MagicMock()
mock_db.commit = MagicMock()
mock_db.refresh = MagicMock()
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=mock_repo,
)
svc = LipsyncService(mock_db, client=mock_mediakit, cosyvoice_service=mock_cosyvoice)
with pytest.raises(MediaKitError) as exc_info:
with pytest.raises(CosyVoiceError, match="TTS 服务不可用"):
svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="longxiaochun_v3",
script_text="测试文本",
)
assert exc_info.value.code == "TTSSynthesisFailed"
# 不应提交到 MediaKit
mock_mediakit.submit_lipsync.assert_not_called()
# 应该记录了失败状态
added_job = mock_db.add.call_args[0][0]
assert added_job.status == "failed"
assert "TTS" in added_job.error_message
def test_create_job_api_failure(self, mock_mediakit, mock_cosyvoice):
"""MediaKit 提交失败."""
@@ -434,165 +399,78 @@ class TestLipsyncServiceUnit:
assert result.status == "completed"
def test_create_job_stores_tts_audio_url(self, mock_mediakit, mock_cosyvoice):
"""v3: TTS 直生模式下 job.audio_url 为转存到自家 OSS 的永久地址."""
"""#1809: 验证 jobaudio_url 来自 TTS 合成结果."""
from app.services.lipsync_service import LipsyncService
mock_db = MagicMock()
mock_repo = MagicMock()
mock_repo.get.return_value = None
mock_db.add = MagicMock()
mock_db.flush = MagicMock()
mock_db.commit = MagicMock()
mock_db.refresh = MagicMock()
with (
patch("app.services.lipsync_service.get_shared_storage_service") as storage_patch,
patch("app.services.lipsync_service.safe_download_bytes") as dl_patch,
):
storage_patch.return_value.upload_file.return_value = "https://my-oss/permanent.mp3"
dl_patch.return_value = b"audio-bytes"
svc = LipsyncService(mock_db, client=mock_mediakit, cosyvoice_service=mock_cosyvoice)
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=mock_repo,
)
job = svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="my-clone-voice",
script_text="这是一段测试文本",
)
# job.audio_url 是转存 OSS 后的永久地址
assert job.audio_url == "https://my-oss/permanent.mp3"
def test_create_job_direct_audio_skips_tts(self, mock_mediakit, mock_cosyvoice):
"""v3: 直接音频模式(传 audio_url)不触发 TTS,原样把 audio_url 提交 MediaKit."""
from app.services.lipsync_service import LipsyncService
mock_db = MagicMock()
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=MagicMock(),
)
job = svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
audio_url="https://example.com/direct-audio.mp3",
voice_id="my-clone-voice",
script_text="这是一段测试文本",
)
mock_cosyvoice.submit_synthesize_task.assert_not_called()
call_kwargs = mock_mediakit.submit_lipsync.call_args
assert call_kwargs.kwargs["audio_url"] == "https://example.com/direct-audio.mp3"
assert job.audio_url == "https://example.com/direct-audio.mp3"
# job.audio_url 应该是 TTS 返回的 URL
assert job.audio_url == "https://oss.example.com/tts-output.mp3"
class TestErrorHandling:
"""v3: 音色解析与错误码在 service 层处理,路由层做 HTTP 状态码映射."""
"""#1809 补充:错误返回 400 而非 500."""
def test_voice_id_resolve_forbidden(self, mock_mediakit, mock_cosyvoice):
"""v3: 克隆音色属于他人时 service._resolve_voice_id 抛 VoiceForbidden(路由映射 403."""
from app.services.lipsync_service import LipsyncService
from app.services.mediakit_client import MediaKitError
def test_voice_id_resolve_failure_returns_400(self, mock_mediakit, mock_cosyvoice):
"""voice_clone_repo 查询异常时返回 400 而非 500."""
from app.api.routes.lipsync import _resolve_voice_id
from fastapi import HTTPException
mock_db = MagicMock()
mock_repo = MagicMock()
other_profile = MagicMock()
other_profile.user_id = "user-other"
other_profile.voice_id = "cv-voice-1"
mock_repo.get.return_value = other_profile
mock_repo.get.side_effect = Exception("DB connection error")
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=mock_repo,
)
with pytest.raises(HTTPException) as exc_info:
_resolve_voice_id("bad-voice-id", "user-1", mock_repo)
assert exc_info.value.status_code == 400
assert "voice_id" in str(exc_info.value.detail)
with pytest.raises(MediaKitError) as exc_info:
svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="clone-profile-id",
script_text="测试",
)
assert exc_info.value.code == "VoiceForbidden"
mock_mediakit.submit_lipsync.assert_not_called()
def test_voice_id_resolve_not_ready(self, mock_mediakit, mock_cosyvoice):
"""v3: 克隆音色尚未生成 voice_id 时抛 VoiceNotReady(路由映射 400."""
def test_create_job_value_error_returns_400(self, mock_mediakit):
"""ValueError(参数无效)返回 400 而非 500."""
from app.services.lipsync_service import LipsyncService
from app.services.mediakit_client import MediaKitError
mock_db = MagicMock()
mock_repo = MagicMock()
profile = MagicMock()
profile.user_id = "user-1"
profile.voice_id = "" # 克隆未完成
mock_repo.get.return_value = profile
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=mock_repo,
)
with pytest.raises(MediaKitError) as exc_info:
svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="clone-profile-id",
script_text="测试",
)
assert exc_info.value.code == "VoiceNotReady"
def test_tts_value_error_mapped_to_invalid_param(self, mock_mediakit):
"""v3: CosyVoice 抛 ValueError(参数无效)被包装为 TTSInvalidParam(路由映射 400."""
from app.services.lipsync_service import LipsyncService
from app.services.mediakit_client import MediaKitError
mock_cosyvoice = MagicMock()
mock_cosyvoice.submit_synthesize_task.side_effect = ValueError("voice_id 为空")
mock_cosyvoice.synthesize_speech.side_effect = ValueError("voice_id 为空")
mock_db = MagicMock()
mock_repo = MagicMock()
mock_repo.get.return_value = None
svc = LipsyncService(mock_db, client=mock_mediakit, cosyvoice_service=mock_cosyvoice)
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=mock_repo,
)
with pytest.raises(MediaKitError) as exc_info:
# Service 层会 catch CosyVoiceError 但 ValueError 会穿透
# 路由层 catch ValueError → 400
with pytest.raises(ValueError):
svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="some-voice",
voice_id="",
script_text="test",
)
assert exc_info.value.code == "TTSInvalidParam"
def test_missing_both_inputs_raises_invalid_input(self, mock_mediakit, mock_cosyvoice):
"""v3: 既无 audio_url 又无 voice_id+script_text 时抛 InvalidInput(路由映射 400."""
def test_create_job_unexpected_exception_returns_400(self, mock_mediakit):
"""未预期的异常应被路由层捕获返回 400 而非 500."""
from app.services.lipsync_service import LipsyncService
from app.services.mediakit_client import MediaKitError
mock_cosyvoice = MagicMock()
mock_cosyvoice.synthesize_speech.side_effect = RuntimeError("unexpected")
mock_db = MagicMock()
svc = LipsyncService(
mock_db,
client=mock_mediakit,
cosyvoice_service=mock_cosyvoice,
voice_clone_repo=MagicMock(),
)
svc = LipsyncService(mock_db, client=mock_mediakit, cosyvoice_service=mock_cosyvoice)
with pytest.raises(MediaKitError) as exc_info:
with pytest.raises(RuntimeError):
svc.create_job(
user_id="user-1",
video_url="https://example.com/video.mp4",
voice_id="test-voice",
script_text="test",
)
assert exc_info.value.code == "InvalidInput"
mock_cosyvoice.submit_synthesize_task.assert_not_called()
mock_mediakit.submit_lipsync.assert_not_called()
-4
View File
@@ -101,7 +101,6 @@ class TestTTSPreviewEndpoint:
text="你好世界",
voice_id="longxiaochun",
speed=1.0,
emotion="",
)
def test_preview_with_speed(self):
@@ -147,7 +146,6 @@ class TestTTSPreviewEndpoint:
text="测试",
voice_id="v1",
speed=1.5,
emotion="",
)
def test_preview_cosyvoice_error_returns_502(self):
@@ -332,7 +330,6 @@ class TestTTSPreviewEndpoint:
text="克隆音色测试",
voice_id="cosyvoice_actual_voice_123",
speed=1.0,
emotion="",
)
# Verify repo was queried with the UUID
mock_clone_repo.get.assert_called_once_with("abc123-uuid-of-profile")
@@ -409,7 +406,6 @@ class TestTTSPreviewEndpoint:
text="预设音色测试",
voice_id="longxiaoxia_v3",
speed=1.0,
emotion="",
)
def test_preview_clone_voice_wrong_user_returns_403(self):