# H3 공유 큐 입력·프롬프트 가이드

이 문서는 `https://h3-shared-queue.pages.dev` 주소만 전달받은 사람이나 AI
에이전트가 어떤 파일과 정보를 어떤 형식으로 넣어야 하는지 설명합니다.
실행 안전 규칙과 전체 명령은 [/AGENTS.md](/AGENTS.md)를 먼저 따르세요.

## 제출 전에 필요한 정보

| 입력 | 이미지 I2V | 영상 R2V |
|---|---|---|
| 주 레퍼런스 | JPG/PNG/WebP 1장, 최대 15MB | MP4/WebM/MOV/M4V 1개, 2–15초, 최대 95MB |
| 추가 레퍼런스 | 없음 | JPG/PNG/WebP 최대 9장, 각 15MB |
| 프롬프트 태그 | 주 이미지는 `<Picture 1>` | 영상은 `<Video 1>`, 추가 이미지는 업로드 순서대로 `<Picture 1>`…`<Picture 9>` |
| 프로필 | 기본 `turbo` | 자동 `ref2va` 20-step |
| 필수 텍스트 | 작업 이름, 전체 영상 프롬프트 | 작업 이름, 각 레퍼런스 역할을 포함한 전체 영상 프롬프트 |
| 선택값 | profile, width, height, frames, seed | width, height, frames, seed |

`H3_SHARED_KEY`는 사람 소유자가 별도 비밀 채널이나 환경변수로 제공해야 합니다.
공개 URL, 문서, 프롬프트 파일에는 키가 들어 있지 않습니다.
영상 제출 컴퓨터에는 길이 검증용 `ffprobe`가 설치되어 있어야 합니다.

## 어떤 모드를 선택할까

- 정지 이미지의 인물·제품·구도를 움직이려면 이미지 I2V를 사용합니다.
- 기존 영상의 동작, 카메라 리듬, 인물 또는 음성 특징을 참조하려면 영상 R2V를
  사용합니다.
- 영상의 움직임은 가져오되 제품·브랜드·의상은 별도 사진으로 고정하려면
  `<Video 1> + <Picture 1>...` 혼합 R2V를 사용합니다.
- 모션 영상과 제품 사진을 한 장의 보드나 하나의 합성 영상으로 합치지 마세요.
  파일을 각각 업로드해야 H3가 역할을 분리해 읽습니다.

## 이미지 I2V 프롬프트 템플릿

프롬프트 본문은 영어로 작성하고 실제 대사만 원하는 언어를 유지하는 편이
안정적입니다. 대괄호 부분을 실제 정보로 바꾸세요.

```text
For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced.

integrated_multimodal_description: [Shot 1] Live-action [phone / observational / product] footage. Preserve the exact [person, face, body proportions, clothing, product geometry, room layout, and lighting] established by <Picture 1>. From 0.00 to [time] seconds, [small anticipatory movement]. From [time] to [time] seconds, [one primary continuous action with clear contact, weight, gravity, and follow-through]. From [time] seconds to the end, [physical result and one restrained reaction]. The camera performs one [static shot / slow push-in / restrained tracking move]. Keep identity, anatomy, object dimensions, background layout, exposure, and light direction stable. Natural real-time pace; no cuts, slow motion, speed ramps, duplicated limbs, face reshaping, or unexplained object changes.

overall_soundscape: [location ambience], [synchronized contact/footstep/fabric/object sounds], [breathing or dialogue]. No subtitles.
non_diegetic_music: [specific restrained music] or N/A
```

한 컷에는 주요 행동 하나와 카메라 움직임 하나만 지정하세요. “cinematic, 8K,
masterpiece” 같은 품질 단어를 반복하기보다 접촉, 무게 이동, 관성, 시선, 빛의
유지를 관찰 가능한 문장으로 적는 편이 낫습니다.

## 영상 R2V + 제품·브랜드 이미지 템플릿

추가 이미지가 없다면 `<Picture N>` 정의를 삭제합니다. 원본 영상의 소리를 실제로
참조할 때만 `<Audio 1>`을 넣으세요.

```text
Source and reference roles:
<Video 1> provides only the shot order, cut timing, camera movement, framing, pose timing, hand motion, group blocking, and pacing. Do not reproduce its performers, faces, wardrobe, title artwork, choreography, brands, music, vocals, or lyrics unless explicitly assigned below.
<Picture 1> is the strict product reference. Preserve its exact geometry, proportions, material, color, labels, logo placement, and handling scale.
<Picture 2> is the strict brand identity reference. Preserve only the specified wordmark, typography, and core palette.
[Define <Picture 3> through <Picture 9> separately when present.]
[<Audio 1> provides only the referenced voice / rhythm / ambience.]

For the target video, at 0.00 seconds into the target video, <Picture 1> (from [Shot 1]) is fully referenced as the invariant product design, <Picture 2> is fully referenced as the invariant brand identity, and the motion structure of <Video 1> begins.

integrated_multimodal_description: [Shot 1] Create one cohesive [duration]-second, vertical 9:16 [campaign / live-action / observational] sequence. Follow the exact motion and camera rhythm of <Video 1>, while replacing the source performers and scene with [target subject and setting]. Use the exact product from <Picture 1> at a physically consistent scale and apply the brand rules from <Picture 2>. Describe the chronological action, contact, result, and restrained ending. Keep faces, anatomy, product geometry, labels, wardrobe, background layout, and lighting stable. Natural real-time pace; no unrequested text, cuts, slow motion, speed ramps, duplicated limbs, morphing, or source-video brands.

overall_soundscape: [new ambience and synchronized physical sounds, or the exact limited role of <Audio 1>]. No subtitles.
non_diegetic_music: [new music relationship] or N/A
```

레퍼런스 역할은 반드시 분리해서 적으세요. 예를 들어 `<Video 1>`에는 동작과
카메라만, `<Picture 1>`에는 제품 형상만, `<Picture 2>`에는 로고와 색상만
할당합니다. 원본 영상의 배우·얼굴·상표를 무심코 복제하지 않도록 가져오지 않을
요소도 명시하는 편이 안전합니다.

## 권장 프로필

| 프로필 | 따뜻한 RTX 5090 예상 | 적합한 작업 |
|---|---:|---|
| `turbo-draft` | 55–80초 | 구도와 동작 후보 비교 |
| `turbo` | 70–105초 | 대부분의 약 13초 게시 후보 |
| `draft` | 100–150초 | 비-Turbo 움직임 비교 |
| `balanced` | 140–210초 | 손, 제품, 복잡한 접촉 |
| `final` | 260–380초 | 선별된 히어로 컷 |
| `native` | 360–540초 | 네이티브 약 1MP가 꼭 필요한 최종 컷 |
| `ref2va` | 240–360초 | 영상 동작·카메라·음성 참조 |

예상 시간은 모델이 이미 로드된 5090 기준이며 실제 시간은 매 결과의 `H3 생성`과
`전체 제작` 필드에 기록됩니다. 이미지 작업은 별도 지정이 없으면 `turbo`를
선택하세요.

## 그대로 실행할 수 있는 제출 예시

먼저 같은 URL에서 클라이언트를 받습니다.

```bash
curl --fail --show-error \
  https://h3-shared-queue.pages.dev/h3-client.py \
  --output h3-client.py
```

이미지 I2V:

```bash
python3 h3-client.py submit /absolute/path/reference.jpg \
  --prompt-file /absolute/path/prompt.txt \
  --name descriptive-output-name \
  --profile turbo \
  --ready \
  --json
```

영상 R2V와 추가 제품·브랜드 이미지:

```bash
python3 h3-client.py submit /absolute/path/motion.mp4 \
  --reference-image /absolute/path/product.jpg \
  --reference-image /absolute/path/brand-board.png \
  --prompt-file /absolute/path/prompt.txt \
  --name mixed-reference-shot \
  --ready \
  --json
```

`--reference-image`의 순서가 `<Picture 1>`, `<Picture 2>` 번호가 됩니다.
`--ready`는 미래 배치의 실행 준비 상태일 뿐 GPU를 켜지 않습니다. 반환된 전체 작업
ID를 사용자에게 알려주세요. 모든 준비 작업을 배치로 묶으라는 명시적 요청을 받았을
때만 `python3 h3-client.py run`을 실행합니다.

## 16초 이상 영상

H3 한 컷은 최대 약 15초입니다. 목표가 16–90초이면 먼저 하나의 작업을 만든 뒤
대시보드의 `16초 이상 멀티컷` 도구를 사용하세요. 약 12초 단위로 2–8개의 편집
가능한 컷이 만들어집니다. 각 컷의 행동과 연결점을 따로 수정한 뒤 준비해야 합니다.
모든 컷 생성이 끝나면 Mac 브리지가 공통 해상도·H.264·AAC로 자동 결합합니다.

인증 API를 직접 쓸 때는 다음 요청입니다.

```http
POST /api/jobs/JOB_ID/cut-plan
Content-Type: application/json

{"target_duration_seconds": 30}
```

## 결과와 후처리

완료 결과에는 순수 H3 생성 시간과 다운로드·업로드를 포함한 전체 제작 시간이
기록됩니다. Mac 브리지는 대표 프레임을 분석해 `바로 사용`, `후처리 권장`,
`재생성 권장` 중 하나와 작업 순서, 복사 가능한 후처리 프롬프트를 제공합니다.
분석을 다시 요청하는 API는 다음과 같습니다.

```http
POST /api/jobs/JOB_ID/results/RESULT_ID/postprocess
```

## 최종 체크리스트

- 주 레퍼런스가 절대 경로이며 형식·크기·영상 길이가 제한 안에 있는가
- 프롬프트가 레퍼런스마다 하나의 명확한 역할을 지정하는가
- 이미지 태그를 `<Image 1>`이 아니라 `<Picture 1>`로 썼는가
- 한 컷에 연속 행동 하나와 카메라 움직임 하나만 있는가
- 유지할 얼굴·제품·라벨·의상·배경·빛을 구체적으로 적었는가
- 공간음, 접촉음, 대사, 음악 유무를 적었는가
- 사용자가 원한 경우에만 `--ready` 또는 배치 `run`을 실행하는가
- `h3-up`, `h3-down`, Vast/RunPod API는 인간 소유자에게 맡겼는가
