ClaudeMods
☰
KO
● 0 명 접속 중 · 조회 0 회
후원프로젝트 제출
GitHub 저장소 · 작성자 zeropage-onryo

zp-ci-watch

git push 또는 PR 병합 후 해당 커밋의 GitHub Actions 실행을 상태 줄과 토스트로 추적하고 초록색 또는 빨간색으로 알립니다. 결과를 확인하려고 세션이 루프에서 계속 잠들어 있을 필요가 없습니다. `/ci`는 한 번만 묻습니다.

zeropage-onryo@zeropage-onryo

zeropage-onryo/video-pipeline/tree/main/.claude/skills/zp-ci-watch

번역 완료

이 mod 소개

ZPF Studio — 아이디어 한 줄에서 완성된 영상까지

아이디어를 하나 건네세요. 선택할 수 있는 장면을 여러 개 만들고, 고른 장면을 디렉터 노트로 다듬어 원하는 상태가 되면 완성된 영상이나 이미지로 렌더링합니다. 탭 6개와 메모 앱을 오갈 필요 없이 6개 생성 플랫폼을 하나의 스튜디오에서 다룹니다.

idea spark
    │
    ▼
several scenes generated off that one idea  ──── you pick one  (pick_rate)
    │
    ▼
director mode — one note at a time: "shot slower", "hold the reveal a beat longer"
    │
    ▼
prompt enhanced (Gemini 3 Flash)
    │
    ├──▶ Nano Banana Pro ──▶ keyframe you approve ──┐
    │                                               │
    └───────────────────────────────────────────────┴──▶ Runway ──▶ finished clip
                                                                        │
                                                                        ▼
                                                          posted ──▶ view counts
                                                                        │
                                                     feeds the next spark ┘

키프레임은 옆으로 빠지는 가지가 아닙니다. 향상된 프롬프트는 두 렌더 노드에 모두 들어가고, 승인한 스틸 이미지는 Runway의 이미지 포트로 들어갑니다. 그래서 클립은 텍스트만이 아니라 이미 좋다고 승인한 이미지에서 시작합니다. 장면 하나는 프롬프트 하나와 샷 하나이며, 렌더 결과 자체가 납품물입니다.

플랫폼마다 방언이 조금씩 다릅니다. Veo가 원하는 것과 Kling이 원하는 것, Seedance가 원하는 것은 각각 다릅니다. 아이디어를 손으로 옮기는 데 시간이 들지만 창작 작업은 아닙니다. 여기서 이를 한데 모읍니다. 통제된 하나의 샷 어휘를 입력하고, 플랫폼에 맞는 프롬프트를 출력하고, 플랫폼별 커넥터로 렌더를 보내며, 모든 시도를 기록합니다. 어떤 프롬프트가 2번 만에 통하고 어떤 프롬프트가 9번 걸리는지 알 수 있습니다.

아이디어는 가장 잘 맞는 곳에서 가져옵니다. Zero Page는 형식 골격, 즉 다른 장면에도 옮겨 쓸 수 있는 구조를 활용하고 실제 성과에 따라 순위를 매깁니다. Antihero는 반복해서 등장하는 스타를 실제로 촬영한 방에 배치합니다. 둘 다 영감 계정, 직접 쓴 글의 검색 라이브러리, 직접 게시한 결과의 근거를 활용할 수 있습니다. 자신의 자료를 근거로 삼는 것은 토글이지, 통과해야 하는 문이 아닙니다.

상태: 출시 전. 로그인, 계정, 기능 게이팅은 작동하지만 아직 다른 사람에게 공개하지 않았습니다. 실제 현재 상태를 참고하세요.

하는 일

https://github.com/user-attachments/assets/967a5d87-b345-404c-8d84-d72477798b9b

Spark → 고를 수 있는 장면. 아이디어와 브랜드를 바탕으로 한 번의 호출에서 여러 장면을 만듭니다. 선택지는 따로 굴린 결과가 아니라 서로 대비되도록 생성됩니다. 각 장면은 카메라, 구도, 움직임, 조명, 다이에제틱 사운드를 포함한 완성된 붙여넣기용 장면 프롬프트입니다. 고른 장면은 그것을 만든 프롬프트의 해시와 함께 pick_rate에 기록됩니다.

Director mode — 장면이 영상이 되는 곳. 두 부분으로 나뉩니다. 대화 쪽은 쉬운 말로 노트 하나를 받습니다. 예를 들어 "shot 2 slower", "hold the reveal a beat longer", "move it to the bedroom"처럼 말하면 Gemini 3 Flash가 저장된 장면을 그 자리에서 고칩니다. 다시 생성하지 않고, 고른 제목과 hook, logline도 건드리지 않습니다. 캔버스 쪽은 렌더링합니다. 샷의 프롬프트를 향상하고, 그 프롬프트를 Nano Banana Pro에 보내 키프레임을 만든 다음, 키프레임을 Runway에 보냅니다. 클립은 텍스트만이 아니라 이미 승인한 스틸에서 시작합니다. 참고 이미지를 붙이고 RAG 검색을 하는 일은 백엔드에서 처리하므로 노드가 더 생기지 않습니다.

노트 하나가 작업을 망치지 않도록 두 가지 보호가 있습니다. 첨부 자료는 수정 뒤에도 남습니다(모델이 새 값을 명시적으로 돌려주지 않는 한 샷은 reference_image와 media_url을 유지하므로, 문구만 바꿔도 이미 렌더한 클립이 분리되지 않습니다). 깨진 출력도 저장되지 않습니다. 해석할 수 없는 응답, 빈 샷 목록, 샷의 절반 넘게 사라진 계획은 오류가 되고 저장된 장면은 그대로 유지됩니다.

어휘 하나, 플랫폼 6개. 각 샷에는 Veo, Kling, Runway, Seedance, LTX, Wan에 맞는 붙여넣기용 프롬프트가 있습니다. 통제된 하나의 샷 어휘로 작성하고 플랫폼마다 렌더러를 하나씩 둡니다. 플랫폼이 지원하면 참고 기반과 Extend Video 시퀀스 연결도 사용합니다.

렌더를 보냅니다. runway.py, veo.py, midjourney.py, nano_banana.py는 같은 형태를 공유하는 게이트된 커넥터입니다. 일일 한도, 호출 전 비용 추정, 모든 시도의 기록, 설정되지 않은 커넥터를 실행 중단 대신 이유와 함께 낮은 기능으로 처리하는 동작이 들어 있습니다. 장면 보드에서 한 번 클릭하면 됩니다.

모든 샷을 생성합니다. 이제 카메라 경로는 없습니다. 샷의 source는 실제 참고 자료가 생성의 기반이 되었는지를 기록합니다. 연기 테이크나 방의 플레이트가 샷의 reference_image로 들어가는 경우입니다. 샷이 파이프라인을 빠져나갔는지를 기록하는 값은 아닙니다. 참고 자료는 강화 요소이지 게이트가 아니며 RAG grounding도 같습니다.

선택에서 배웁니다. 계획한 콘셉트와 실제로 만든 콘셉트를 생성한 프롬프트의 해시와 함께 기록합니다. 프롬프트가 바뀌었을 때 논쟁하는 대신 측정할 수 있습니다. 게시한 영상과 조회 수도 다시 들어오며 같은 경과 시간 기준으로 비교하므로, 1년 된 영상이 누적 합계만으로 이길 수 없습니다. rework.py는 처음부터 다시 시작하지 않고 이 근거에서 다음 목록을 구상합니다.

브랜드별로 분리됩니다. 영감 경로를 브랜드별로 나누어 근거가 서로 새지 않습니다.

스튜디오

/ui는 브라우저에서 쓰는 제품입니다. FastAPI + Jinja2로 동작하며 빌드 단계가 없습니다.

  • Composer — 첫 화면에서 참고 자료, 방, 캐릭터, 소품을 올리고 선택한 뒤 생성합니다.
  • Scene board — 장면을 만들고 샷별 프롬프트를 복사하고 렌더를 시작한 뒤 클립을 다시 붙입니다.
  • Assets gallery — 날짜별로 묶인 미디어 패널입니다.
  • Holds — 보류된 실행마다 사람이 읽을 수 있는 이유와 approved / rejected 상태가 있습니다.
  • Jobs — 오래 걸리는 작업을 폴링하지 않고 SSE로 스트리밍합니다.
  • Sign-in — Supabase Auth(Google, Discord, email/password)를 사용하고 서버에서 PyJWT로 확인합니다. 설정된 제공자만 모달에 표시하므로 클라이언트 시크릿이 없으면 버튼을 숨길 뿐 페이지를 깨뜨리지 않습니다.

그 뒤의 JSON API는 기능으로 게이트되며 실제 키가 있는지에서 실시간으로 파생됩니다. 정적인 dict가 아닙니다. 뒤의 엔드포인트가 실제로 실행될 수 있을 때만 컨트롤이 나타납니다.

전체를 세우는 규칙

점수 없이 크레딧을 쓰지 않습니다. 생성된 프롬프트마다 두 단계 심사를 거칩니다. 먼저 비용 없는 구조 검사, 다음으로 닫힌 상태에서 실패하는 엄격한 LLM 기준입니다. 읽을 수 없는 판정은 기본 통과하지 않고 0점이 됩니다. 한 번의 실행에 있는 모든 프롬프트가 통과해야 하며, 그렇지 않으면 실행이 보류되고 일부 렌더만 진행되지 않습니다.

깨뜨리지 말고 낮춰서 계속합니다. Postgres가 없거나, 커넥터가 설정되지 않았거나, 설명한 장소가 없거나, 기능이 아직 들어오지 않았어도 실행은 계속되고 이유를 말합니다. 일부러 둔 예외는 promptgen과 locations입니다. 이곳에서는 모델 호출 자체가 납품물이며 그 위에 얹힌 기록 작업이 아니므로 크게 실패합니다.

프롬프트는 요청하고 코드는 강제합니다. 모델 출력을 설명한 장소, 샷 어휘, 도구 레지스트리와 대조합니다. 불일치는 모두 저장된 결과에 보이는 경고로 나타나며 거부되지는 않습니다. 모델은 방과 어휘를 환각할 수 있으므로 결정하는 사람이 그 사실을 봐야 합니다.

실행해서 확인합니다. 이 프로젝트의 실제 결함은 코드 리뷰도 통과했고 자체 테스트도 통과했습니다. 서버를 시작해 클릭하거나 테스트 모음이 조용히 느려졌다는 사실을 알아차렸을 때 발견되었습니다.

품질 평가

네 개의 스코어러를 일부러 분리합니다.

| Module | Judges | |---|---| | prompt gate (orchestrator) | 이 프롬프트는 크레딧을 쓸 만큼 잘 구성되었나? Subject / camera / motion / lighting / coherence를 각각 0–2로 평가하고 PROMPT_GATE_MIN과 비교합니다. 닫힌 상태에서 실패합니다. | | uncanny_judge.py | 브랜드에 맞는지 보는 게이트입니다. 고정된 기준을 쓰므로 배울 기록이 없던 첫날부터 작동합니다. 채널을 자동 운용해도 안전하게 만드는 부분입니다. | | taste_judge.py | 이 크리에이터 자신의 기록을 기준으로 평가합니다. /holds에서 승인하거나 거절한 것, 잘됐다고 표시한 것, 성공한 게시물과 실패한 게시물의 특징을 보고 "they'll like this."를 예측합니다. | | quality.py | 파이프라인이 이미 사용하는 같은 Gemini 모델과 DeepEval의 LLM judge 지표로 검색된 문맥에 대한 충실도와 답변 관련성을 평가합니다. |

프롬프트는 구조적으로 훌륭하지만 톤이 틀릴 수도 있고, 브랜드에는 맞지만 전반적으로 나쁠 수도 있기 때문에 분리했습니다. 점수 하나로 합치면 실패 원인을 알 수 없습니다. evalstore.py가 저장하므로 품질을 느낌이 아니라 추세로 볼 수 있습니다.

실행하기

venv/bin/pip install -r requirements.txt && venv/bin/pip install -e .
venv/bin/uvicorn app.main:app --reload   # the Studio, in the browser

.env에 GEMINI_API_KEY가 필요합니다. 선택 사항으로 YOUTUBE_API_KEY(공개 조회 수, 채널 가져오기), 로그인용 OAuth 클라이언트 자격 증명, 플랫폼별 렌더 키가 있습니다. 없는 키는 시작을 망가뜨리지 않고 이유를 표시하며 해당 기능을 끕니다.

각 단계에는 CLI도 있습니다. python -m src.locations, src.shootgen, src.promptgen, src.director, src.genlog, src.orchestrator, src.trigger, src.rework, src.autopilot, src.scheduling, src.accounts입니다. 전체 목록은 CLAUDE.md를 보세요.

참고 라이브러리(RAG)

글이 학습할 텍스트, 즉 브랜드 노트, 과거 스크립트, 좋아하는 영화에 대한 메모, 플랫폼 프롬프트 참고 자료를 잘라 gemini-embedding-001로 임베딩하고 PostgreSQL + pgvector에 저장합니다. 문서와 쿼리는 모델이 비대칭이므로 서로 다른 작업 유형으로 임베딩합니다. 생성 시 spark, brand, mood가 쿼리가 되고 가장 가까운 청크가 어조와 구조 참고로 주입됩니다. 검색은 CRAG로 평가됩니다(crag.py). 첫 결과가 약하면 조용히 쓰지 않고 쿼리를 한 번 다시 씁니다. Postgres가 없어도 grounding 없이 실행을 계속하고 그 사실을 알립니다.

운영 환경의 모든 CRAG 결정은 참고 텍스트를 저장하지 않은 채 기록됩니다. 초기 및 재시도 점수, 재시도를 실행했는지, 점수가 나아졌는지, 채택했는지, 당시의 임계값과 참고 라이브러리 지문을 남깁니다. 비공개 /studio Stats 보기에는 제품 텔레메트리가 표시됩니다. /ui도 같은 검색 경로를 쓰지만 내부 진단을 노출하지 않습니다.

성공한 Nano Banana 이미지와 Runway 비디오는 정확한 프롬프트, provider/model, 미디어 형식, 사용 가능한 장면 메타데이터와 함께 /ui의 Asset Bank에도 게시됩니다. Nano 이미지는 나중의 이미지 참고로 계속 선택할 수 있고 Runway 클립은 갤러리에 나타나지만 이미지 전용 선택기에서는 제외됩니다. 프롬프트와 모델 설명은 RAG assets 선반에 색인되며, 일시적인 벡터 저장소 장애가 완료된 렌더나 로컬 Asset Bank 기록을 버리지는 않습니다.

docker compose up -d                       # Postgres + pgvector

venv/bin/python -m src.rag ingest .claude/skills/video-prompting/references/models/veo3/prompting.md --domain ai_prompting
venv/bin/python -m src.rag query "stillness broken once" --k 5
venv/bin/python -m src.shootgen --spark "gearing up ritual"   # picks up references on its own

venv/bin/python -m src.rag_eval eval_cases.json --k 5         # hit@k, MRR against labeled cases

Dev Studio 평가 실행은 같은 golden 질문을 두 번 사용합니다. 한 번은 one-shot 검색으로, 한 번은 완전한 CRAG 재작성 경로로 실행합니다. 기본 결과와 CRAG Hit@k/MRR, 재쿼리 비율, 재시도 점수가 개선되는 빈도, 재시도가 채택되는 빈도, 재시도가 실제로 사람이 라벨링한 예상 소스를 돌려주는지를 보고합니다. 재쿼리 비율이 낮다고 그 자체로 성공으로 보지 않습니다. 검색 정확도를 유지하거나 높여야 합니다. 실행은 쿼리 세트 지문, 임베딩/재작성 모델, 임계값, 참고 라이브러리 수와 지문을 기록해 서로 다른 구성을 동등한 비교로 제시하지 않습니다.

모든 청크에는 필수 domain 선반 라벨이 있으므로 쿼리에서 의미 범위를 제한할 수 있습니다.

설치

먼저 작성자의 README에서 marketplace와 플러그인 이름을 확인하세요. 저장소 구조에 따라 명령어가 달라질 수 있습니다.

claude plugin marketplace add zeropage-onryo/video-pipeline
claude plugin install zp-ci-watch
원문 / README

ZPF Studio — an idea spark to a finished video

Give it an idea. It generates several scenes to choose between, you shape the one you picked with director notes until it's right, and it renders a finished video or image. One studio across six generation platforms instead of six tabs and a notes app.

idea spark
    │
    ▼
several scenes generated off that one idea  ──── you pick one  (pick_rate)
    │
    ▼
director mode — one note at a time: "shot slower", "hold the reveal a beat longer"
    │
    ▼
prompt enhanced (Gemini 3 Flash)
    │
    ├──▶ Nano Banana Pro ──▶ keyframe you approve ──┐
    │                                               │
    └───────────────────────────────────────────────┴──▶ Runway ──▶ finished clip
                                                                        │
                                                                        ▼
                                                          posted ──▶ view counts
                                                                        │
                                                     feeds the next spark ┘

The keyframe is not a side branch: the enhanced prompt feeds both render nodes, and the approved still feeds Runway's image port, so the clip starts from an image you already said yes to rather than from text alone. A scene is one prompt and one shot — the render is the deliverable.

Every platform speaks a slightly different dialect — Veo wants one thing, Kling another, Seedance another. Moving an idea between them by hand is where the time goes, and none of it is creative work. This aggregates that: one controlled shot vocabulary in, platform-native prompts out, renders dispatched through per-platform connectors, and every attempt logged so you learn which prompts land in two tries and which take nine.

Ideas come from wherever they work best. Zero Page rides format skeletons — the structure that travels — ranked by what's actually performing. Antihero grounds in real photographed rooms with a recurring star. Either can draw on inspiration accounts, a retrieval library of your own writing, or evidence from your own posted results. Grounding in your own material is a toggle, not a gate.

Status: pre-launch. Sign-in, accounts and capability gating work; it isn't open to other people yet. See Where this actually is.

What it does

https://github.com/user-attachments/assets/967a5d87-b345-404c-8d84-d72477798b9b

Spark → scenes you choose between. From an idea and a brand, several scenes are generated in a single call — so the options are varied against each other rather than rolled independently — each one a complete, paste-ready scene prompt with camera, framing, movement, lighting and diegetic sound. Which one you pick is recorded (pick_rate), against the hash of the prompt that wrote it.

Director mode — where a scene becomes video. Two halves. The conversational half takes one note in plain language — "shot 2 slower", "hold the reveal a beat longer", "move it to the bedroom" — and revises the stored scene in place via Gemini 3 Flash, never regenerating it and never touching the picked title, hook or logline. The canvas half renders it: the shot's prompt is enhanced, that enhanced prompt feeds Nano Banana Pro for a keyframe, and the keyframe feeds Runway so the clip starts from a still you already approved rather than from text alone. The reference image and RAG retrieval ride on the backend rather than as extra nodes.

Two protections keep a note from destroying work: attached material survives a revision (a shot keeps its reference_image and media_url unless the model explicitly returns new ones, so a wording change can't detach a clip you already rendered), and broken output never lands — an unparseable response, an empty shot list, or a plan that lost more than half its shots is an error, and the stored scene stays exactly as it was.

One vocabulary, six platforms. Each shot carries a paste-ready, platform-native prompt — Veo, Kling, Runway, Seedance, LTX, Wan — written from one controlled shot vocabulary, one renderer per platform. Reference grounding and Extend Video sequence-chaining where the platform supports them.

Renders, dispatched. runway.py, veo.py, midjourney.py and nano_banana.py are gated connectors sharing one shape: a daily cap, a cost estimate before the call, every attempt logged, and a connector that isn't configured degrading with a stated reason rather than failing the run. One click from the scene board.

Every shot is generated. There is no camera path any more. A shot's source records whether real reference material anchors the generation — an acting take, a room plate, feeding the shot's reference_image — not whether the shot escapes the pipeline. A reference is an enhancement, never a gate, the same way RAG grounding is.

It learns from your choices. Which concepts you plan and which you actually make, each recorded against the hash of the prompt that produced it — so a prompt change becomes something you measure rather than argue about. Posted videos and their view counts feed back in, compared at equal age so a year-old video can't win on accumulated totals. rework.py then ideates the next slate from that evidence rather than from scratch.

Brand-scoped. Inspiration lanes are scoped per brand so grounding never leaks across them.

The Studio

/ui is the product in a browser — FastAPI + Jinja2, no build step.

  • Composer — upload references, pick rooms, characters and props from the front page, generate.
  • Scene board — generate a scene, copy per-shot prompts, fire a render, attach clips back.
  • Assets gallery — date-grouped media panel.
  • Holds — every parked run with a human-readable reason, graded approved / rejected.
  • Jobs — long-running work streamed over SSE rather than polled.
  • Sign-in — Supabase Auth (Google, Discord, email/password), verified server-side with PyJWT. The modal renders only the providers actually configured, so a missing client secret hides a button instead of breaking the page.

The JSON API behind it is capability-gated and derived live from real key presence, never a static dict — a control only appears if the endpoint behind it can actually run.

The rules the whole thing is built on

Nothing spends a credit unscored. Every generated prompt runs a two-layer judge: a zero-cost structural check, then a strict LLM rubric that fails closed, so an unreadable verdict scores 0 rather than passing by default. All prompts in a run must pass or the run holds — no partial renders.

Degrade, don't break. A missing Postgres, an unconfigured connector, an absent described location, a feature that hasn't landed — each degrades to a run that continues and says so. The exceptions are deliberate: promptgen and locations fail loudly, because there the model call is the deliverable rather than bookkeeping on top of one.

Prompts request, code enforces. Model output is checked against reality — described locations, the shot vocabulary, the tool registry — and every mismatch surfaces as a visible warning on a saved result. Nothing is rejected. Models hallucinate rooms and vocabularies, and the human deciding needs to see that.

Verify by running it. Every real defect this project has had passed code review and passed its own tests. They were caught by starting the server, clicking the thing, or noticing the suite had quietly gotten slower.

Judging quality

Four scorers, deliberately separate:

| Module | Judges | |---|---| | prompt gate (orchestrator) | Is this prompt well-formed enough to spend a credit on? Subject / camera / motion / lighting / coherence, 0–2 each, against PROMPT_GATE_MIN. Fails closed. | | uncanny_judge.py | The on-brand gate. A fixed rubric, so it works from day one — before there's any history to learn from. This is what would make a channel safe on autopilot. | | taste_judge.py | Scores against this creator's own record — what they approved and rejected on /holds, what they marked worked, the traits of their winning versus losing posts. Predicts "they'll like this." | | quality.py | Faithfulness and answer relevancy over retrieved context, via DeepEval's LLM-judge metrics on the same Gemini models the pipeline already uses. |

Kept separate because a prompt can be structurally excellent and tonally wrong, or perfectly on-brand and generically bad. One collapsed score makes failures non-diagnostic. evalstore.py persists them, so quality is a trend rather than a vibe.

Running it

venv/bin/pip install -r requirements.txt && venv/bin/pip install -e .
venv/bin/uvicorn app.main:app --reload   # the Studio, in the browser

Needs GEMINI_API_KEY in .env. Optional: YOUTUBE_API_KEY (public view counts, channel import), OAuth client credentials for sign-in, and per-platform render keys. Each absent key disables its feature with a stated reason rather than breaking startup.

Every step also has a CLI — python -m src.locations, src.shootgen, src.promptgen, src.director, src.genlog, src.orchestrator, src.trigger, src.rework, src.autopilot, src.scheduling, src.accounts. See CLAUDE.md for the full list.

Reference library (RAG)

Text you want the writing to learn from — brand notes, past scripts, films-you-admire notes, platform prompting references — chunked, embedded with gemini-embedding-001 (documents and queries embedded with different task types, because the model is asymmetric), and stored in PostgreSQL + pgvector. At generation time the spark, brand and mood become the query, and the closest chunks are injected as tone and structure references. Retrieval is CRAG-graded (crag.py): a weak first pass earns one query rewrite rather than being silently used. No Postgres? The run continues ungrounded and says so.

Every production CRAG decision is logged without storing reference text: initial and retry scores, whether a retry ran, whether it improved the score, whether it was adopted, and the threshold plus reference-library fingerprint used at the time. The private /studio Stats view shows that product telemetry; /ui uses the same retrieval path but never exposes the internal diagnostics.

Every successful Nano Banana image and Runway video is also published to /ui's Asset Bank with its exact prompt, provider/model, media type, and available scene metadata. Nano images remain selectable as later image references; Runway clips appear in the gallery but stay out of image-only pickers. The prompt and model description are indexed on the RAG assets shelf, while a temporary vector-store failure never discards the completed render or its local Asset Bank record.

docker compose up -d                       # Postgres + pgvector

venv/bin/python -m src.rag ingest .claude/skills/video-prompting/references/models/veo3/prompting.md --domain ai_prompting
venv/bin/python -m src.rag query "stillness broken once" --k 5
venv/bin/python -m src.shootgen --spark "gearing up ritual"   # picks up references on its own

venv/bin/python -m src.rag_eval eval_cases.json --k 5         # hit@k, MRR against labeled cases

The Dev Studio eval run uses the same golden questions twice: once with one-shot retrieval and once through the complete CRAG rewrite path. It reports base versus CRAG Hit@k/MRR, re-query rate, how often retry scores improve, how often retries are adopted, and whether retries actually return a human-labelled expected source. A lower re-query rate is not treated as success by itself: it must hold or improve retrieval accuracy. Runs record the query-set fingerprint, embedding/rewrite model, threshold, and reference-library count/fingerprint so unlike configurations are not presented as equivalent comparisons.

Every chunk carries a required domain shelf label, so queries can scope semantically and by hard SQL filter in one pass. Re-ingesting a source replaces its chunks, keyed by path relative to the project root — not basename, or editing/notes.txt and lighting/notes.txt would delete each other.

Tested in CI

Every push runs pytest and ruff, then an eval gate that stands up an ephemeral Postgres + pgvector, re-ingests the library, and runs two gates kept separate so each failure is diagnostic:

  • Retrieval regression — hit@5 and MRR against floors.
  • Generation quality — a 14-case golden set scored on faithfulness, answer relevancy, contextual precision and recall, each against an absolute floor and a regression band versus recorded baselines.

tests/conftest.py blocks all network access during tests, because the same bug landed four times: a test patches one generator, the route changes to call a different one, the patch silently misses, and a real billed API call happens while the test still passes.

The run history is public in the Actions tab, failures included — a judge-model timeout that clipped three golden cases, a deepeval pin after 4.1.10 dropped Python 3.9, a CI-only failure from an unstubbed client.

Tech stack

Python · FastAPI + Jinja2 · SQLite (spaces, concepts, scenes, videos, metric snapshots, generation attempts) · PostgreSQL + pgvector (retrieval library) · Google Gemini — vision, structured generation, embeddings, Gemini 3 Flash for enhancement and director notes, Nano Banana Pro (gemini-3-pro-image-preview) for keyframes · LangGraph + LangSmith (orchestration and tracing) · DeepEval (judge metrics) · Runway / Veo / Midjourney connectors · Instagram + YouTube metrics · Cloudflare R2 · ffprobe · Supabase Auth · pytest + ruff in CI · Docker Compose

Where this actually is

Being straight about the state, because the code will tell you anyway:

  • Not launched. Sign-in and capability gating work, but a fresh signup gets zero membership rows and sees "no account access yet" — membership is granted by hand for v1. Nobody outside has an account.
  • No tenancy yet. Owned tables have no account_id, and the render caps are global rather than per-account. That's the actual blocker to a pilot, scoped in docs/tasks/task-account-tenancy.md.
  • The posting line is deliberately stubbed. generate_render only calls a real renderer when explicitly enabled; no posting API is wired, so even a channel set to auto parks with an explicit reason rather than pretending to post. Instagram and TikTok stats stay manual until developer approvals land.
  • A scene is the unit, not a cut. One scene is one prompt and one shot, so its render is a finished deliverable. Stringing several scenes into a longer edit is still done by hand in Resolve — the timeline assembly was removed deliberately, see the Decisions Log.
  • The measurement loop is structurally complete and statistically empty. The rates and signals are correct and currently meaningless; they need weeks of real posting before a prompt change can be measured rather than argued about.

Roadmap

  • Account tenancy, then a closed pilot — docs/tasks/task-account-tenancy.md
  • Open sign-ups
  • The tool scoreboard — which platform lands which shot type, once enough attempts are logged
  • Verify per-tool camera vocabulary against each platform's current prompt guide
  • Case-study writeup with a demo video

Architecture

Two paths, deliberately different. The request path is what runs when you press Create — a straight line, because a person is standing there waiting for it. The autonomous path is the unattended nightly run, with gates the request path doesn't need.

Both are documented in docs/ARCHITECTURE.md, along with the two known divergences between them.

Decisions Log

2026-07-08 — Two-stage generation (pitch, then edit). src/pitch.py generates a cheap slate of 10 story descriptions from the manifest; a human selects a few; only those get full edit specs (clip in/out points, grade notes, sound notes) generated by src/editgen.py. Selection is the one decision kept manual — everything before and after it is automated.

2026-07-08 — Pipeline operates on proxies, not camera originals. All ingest and analysis runs on DaVinci Resolve proxy files rather than 6K Blackmagic RAW. Reasons: FFmpeg/Whisper can't read .braw natively, processing time scales with file size, and transcription/tagging don't need image quality. Camera originals are only touched by Resolve at final export. Tradeoff: pipeline outputs reference proxy filenames, so filenames must stay consistent between the proxy folder and Resolve's media pool.

2026-07-08 — Footage-first generation instead of script-first. Edit concepts are generated FROM the manifest of real footage, rather than writing scripts and then searching for matching clips. Rejected script-first because it produces edit lists calling for shots that were never filmed. Cost: creative range is bounded by the current library — mitigated by having the model flag footage gaps, which become the next shoot's shot list.

2026-07-08 — Manifest as the interface between stages. The pipeline's stages (ingest → story generation → timeline build) communicate through manifest.json / concepts.json files rather than direct coupling. Reasons: each stage can be run, tested, and debugged independently, and intermediate outputs are human-readable. Tradeoff: filenames act as IDs across stages, so renaming clips after ingest breaks the chain — filenames are treated as immutable once ingested.

2026-07-08 — Gemini API for story generation. Using Google's Gemini rather than other LLM APIs for concept generation. Also relevant: Gemini's native video-input support leaves a clean upgrade path from transcript-based tagging to true visual analysis of clips without changing providers.

2026-07-08 — Model output validated in code, not trusted from the prompt. storygen.py independently verifies that every clip filename in a generated concept exists in the manifest and that in/out points fall within real clip durations, rejecting concepts that fail. The prompt also instructs this, but prompt instructions alone don't prevent hallucinated filenames — validation is enforced at the code layer. Lesson: prompts request, code enforces.

2026-07-08 — Prompts and creative brief live in editable text files. The storygen prompt and brand brief are stored in prompts/ as plain text with placeholder injection, not hardcoded in Python. Prompt and brief tuning is the highest-frequency change in this system; editing text files keeps iteration fast and keeps creative direction separate from logic. Tradeoff: one more layer of files to keep in sync with the code that loads them.

2026-07-08 — Rough cuts target Resolve Studio's scripting API. Timelines are built directly inside the open Resolve project via DaVinciResolveScript, rather than rendering standalone preview files. Edits appear in the real project with media already linked — no export/import round-trip. Fallback: FFmpeg-rendered preview .mp4s if the API path fails or for quick triage.

2026-07-29 — Removed beat-synced cutting. Cut transitions were being snapped to detected or synthetic musical beats. Removed because the timing of a cut is a creative decision and snapping overrode it mechanically — a cut landing on a beat is not the same as a cut landing where the shot wants to end. Also removed librosa and its dependency chain (numba, scipy, soundfile, audioread). Cut points now come only from the clip's own described beats, which is what the edit prompt was already reasoning about.

2026-07-29 — Removed automatic color grading. apply_grade.py applied a saved .drx to every clip on named timelines. Removed because grading is shot-by-shot work and a blanket application produced something that always needed redoing by hand. The grade preset stays in grades/ and is applied manually in Resolve.

2026-07-29 — Superseded 2026-07-08 "Rough cuts target Resolve Studio's scripting API". Removed the Resolve integration entirely. build_timeline.py, apply_grade.py and resolve_edit.py are gone. The pipeline now ends at a validated cut list in concepts.json, which is executed by hand. Reason: assembling the timeline was the least valuable step and the most brittle — it required Resolve running with a project open, matched clips by filename across two systems, and produced an assembly that was always re-cut anyway. The judgment worth automating is which shots and which moments, not the mechanical assembly. Cost: no more one-command rough cut. Accepted, because the rough cut was never the output that got used.

<!-- DRAFT — per RUNBOOK, Decisions Log entries are yours to write. Edit the reasoning (especially the "why now" in the first entry — that's your call, not mine), delete this comment, then commit. -->

2026-07-30 — Added a pre-production phase: locations, then concepts. The pipeline was footage-first end to end — it could only reason about clips that already existed. That meant the hardest part of a one-person operation, deciding what to shoot at all, happened entirely outside the tool. Now locations.py photographs and describes the spaces available (geometry, light sources, textures, workable angles, and what each space won't allow), and shootgen.py generates concepts and ≤6-shot lists grounded in those real rooms. Ported from two React generators that ran as Claude artifacts; moving them in swapped the Anthropic API for this project's existing Gemini client and browser storage for SQLite, which is what makes concepts queryable and comparable rather than trapped in one browser session. Cost: a second meaning for the word "concept" in this repo — concepts.json is a cut list for footage you have, shoot_concepts is a shot list for footage you need. Accepted, with the names kept deliberately distinct.

2026-07-30 — Concepts are grounded in photographed spaces, not imagined ones. validate_concept rejects any shot whose location isn't a space that has actually been photographed and described. This is the pre-production version of the footage-first rule: the same reason edit specs are validated against the manifest applies one step earlier, because a concept set in a room you don't have is worse than no concept — it reads as usable and wastes a shoot day. Tradeoff: you can't generate anything until at least one space is described, and the UI hides the generate button until then rather than offering something that would fail.

2026-07-30 — Recording which concepts actually get shot. shoot_concepts.shot_done, alongside the prompt's hash. Same reasoning as recording which pitches get picked: the decision is already being made, it's free to store, and without it a prompt rewrite can only be argued about rather than measured. Generating ten concepts and shooting one is a different outcome from generating three and shooting two, and shoot_rate() makes that difference visible per prompt version.

2026-08-04 — The pivot: from grounded validator to autonomous content machine. Retired the "grounded inverse of Google Flow" identity — a tool whose selling point was rejecting model output that didn't match reality. The mission now is an automated production pipeline mixing real footage and AI that runs more of itself over time: L1 assisted → L2 grounded generation + measurement → L3 self-improving ideation from performance data → L4 supervised autonomous generate-and-post (gated, dry-run, default off). What changed in code: the 6-shot cap, the one-AI-shot-per-concept slot, and the one-generative-clip-per-edit cap are gone — real and AI shots are co-inputs, each shot carrying source: CAMERA | AI; the AI platform set became data (shot.PLATFORMS, now also Seedance 2.0, LTX-2, Wan 2.2); and a missing described location degrades to an ungrounded run with a note instead of raising. What deliberately did not change: grounding itself (rooms, footage, and the reference library still shape every generation), every human pick recorded against its prompt hash, degrade-don't-break, and the 2026-07-29 removal of the Resolve integration — editing stays manual (an explicit L1 hold). Validators survive as advisories: visible warnings, never gates. Supersedes the rejection clause of "Concepts are grounded in photographed spaces" (2026-07-30) and BUILD_SPEC's "one generated clip per edit, maximum."

2026-08-11 — Documented the LangGraph orchestrator as a first-class architecture section. The graph in src/orchestrator.py was previously explained only in CLAUDE.md, written for an assisting coding agent rather than a reader of this repo. Added an "Architecture — the LangGraph orchestrator" section to this README plus a rendered node/edge diagram at docs/architecture.html, describing the grounding-and-retry loop, the two-layer prompt credit gate, the deliberately stubbed posting line, and the shared hold sink. Reason: the graph's shape is the autonomy story — how much currently runs unattended versus what's still gated — and that was previously implicit in code rather than legible on its own. Nothing in the graph itself changed.

2026-08-11 — Synced this README to the post-production removal. CLAUDE.md already documented that post-production (ingest → pitches → cut lists) was cut from the product in August 2026, but this README's intro, "What it does," Pipeline diagram, Running-it command list, and Reference-library section still described it as live — src.ingest/src.pitch/src.editgen commands that no longer exist, a two-column pre/post-production diagram, and RAG examples grounding "pitches" instead of concepts. Rewrote all of it to the current single-phase pipeline: photograph a space, generate a grounded concept and shot list, shoot it, post it, feed the metrics back. The edit still happens by hand in Resolve — that was already true, just no longer framed as one stage of two. No code changed; this brings the docs in line with what Aug 2026 already did.

2026-08-11 — Fixed a CI-only failure in test_judge_parses_fenced_json_and_clamps_dims. The test mocked generate_with_retry but not _client, so _judge_prompt built a real genai.Client() before ever reaching the mock. That construction is harmless locally (.env has a real GEMINI_API_KEY) but raised in CI's test job, which deliberately has no key — only eval-gate gets one, gated behind a real secret. _judge_prompt's fail-closed except swallowed the exception and returned dims={}, which the test's verdict["dims"]["camera"] lookup then hit as a KeyError. Every other orchestrator test avoids this by stubbing _client through the shared tmp_db fixture; this one test skipped that fixture entirely and needed its own stub. Passed locally, failed in CI — a genuine "trust the log, not the assumption" case, same spirit as the Verify-by-running-it rule above. No production code changed, only the test.

2026-08-20 — Every shot is generated; the camera path is gone. The pipeline was built on real and AI shots as co-inputs, each shot carrying source: CAMERA | AI to say which one it was. That mix is retired. Every shot is now AI-generated, and source was repurposed rather than removed: it records whether real reference material anchors the generation — an acting take, a room plate, feeding the shot's reference_image — not whether the shot escapes the pipeline. A reference is an enhancement, never a gate, exactly as RAG grounding is. Reason: the mix was the last thing forcing the product to be about one operator with two cameras and a house. Cutting it is what let the identity become an aggregator for generating concepts and video, which is the shape someone other than me could use. Cost: the "real material grounds everything, AI extends it" story, which was distinctive and is now simply untrue. Accepted. Direct consequence: Director mode grew a rendering half — enhance the shot prompt with Gemini 3 Flash, generate a Nano Banana Pro keyframe, feed that keyframe to Runway so the clip starts from an approved still instead of from text.

2026-08-27 — From "a solo filmmaker's grounded pipeline" to a multi-platform generation studio. This README described a one-operator, CLI-first pre-production tool whose central rule was that everything is grounded in photographed rooms. Twenty commits since 2026-08-11 made that wrong in both halves. What shipped meanwhile: the ZPF Studio UI (/ui, capability-gated JSON API, jobs over SSE, an eval store), real sign-in (Supabase Auth: Google, Discord and email/password) and an accounts model, gated render connectors for Runway, Veo, Midjourney and Nano Banana, the scene board and director mode, brand-scoped inspiration lanes, and opt-in asset grounding — which demoted grounding from the thesis of the product to a toggle on it. The clearest evidence the old framing had expired is format_feed.py, which states outright that Zero Page rides format skeletons, not rooms. The identity that replaced it: an aggregator for generating content concepts and videos across platforms, where grounding in your own material is one available input rather than the precondition. Also newly documented: the four separate scorers (prompt gate, uncanny_judge's fixed on-brand rubric, taste_judge's learned-from-history rubric, quality.py's DeepEval metrics), which were doing significant work while going unmentioned entirely. Cost: the grounding story was the sharpest thing about the old README, and the new framing is broader and therefore less pointed. Accepted, because a sharp description of the wrong product is worse than an accurate description of the right one. No code changed.

2026-08-27 — Recorded the tenancy gap rather than quietly shipping around it. Adding sign-in made it read as though the product were multi-user. It isn't: list_concepts and get_concept carry no owner predicate, shoot_concepts has no account_id, and the render caps count globally rather than per account — so a second user would see the first's work and exhaust their daily budget. Written up as docs/tasks/task-account-tenancy.md and stated plainly in the README's "Where this actually is" instead of being left for a reader to discover. Reason: the repo is public and linked in job applications; an OAuth flow and an SEO module imply a live product, and being first to say it isn't costs nothing while buying credibility for everything else on the page. Same instinct as logging removals — the honest state of a thing is more useful than the flattering one.

비슷한 프로젝트