From 26a03b87ed5f8f5f40ae6049c8ccc2fa7eb46c19 Mon Sep 17 00:00:00 2001 From: guillergs-ai Date: Tue, 18 Aug 2026 17:08:29 +0200 Subject: [PATCH] docs: add seedance 2.5 prompt contract and technique reference --- .agents/skills/seedance-2-5/SKILL.md | 211 +++++++++++++++++- .../seedance-2-5/reference/techniques.md | 197 ++++++++++++++++ skills/INDEX.md | 2 +- 3 files changed, 408 insertions(+), 2 deletions(-) create mode 100644 .agents/skills/seedance-2-5/reference/techniques.md diff --git a/.agents/skills/seedance-2-5/SKILL.md b/.agents/skills/seedance-2-5/SKILL.md index 282b0911..9c67f4be 100644 --- a/.agents/skills/seedance-2-5/SKILL.md +++ b/.agents/skills/seedance-2-5/SKILL.md @@ -1,7 +1,7 @@ --- name: seedance-2-5 description: | - Generate 4-30 second cinematic video with ByteDance Seedance 2.5 through fal.ai, Volcengine Ark, Runway, or ComfyUI Partner Nodes. Use for long single generations, synchronized audio, and large multimodal reference sets (up to 30 images, 10 videos, and 10 audio clips). + Generate 4-30 second cinematic video with ByteDance Seedance 2.5 through fal.ai, Volcengine Ark, Runway, or ComfyUI Partner Nodes. Use for long single generations, synchronized audio, and large multimodal reference sets (up to 30 images, 10 videos, and 10 audio clips). Also covers the 2.5 prompt contract: section order, multi-shot "Hard cut" breakdown, named locks for continuity across cuts, the asset and character-sheet method, voice conditioning, and iteration discipline. --- # Seedance 2.5 @@ -22,6 +22,11 @@ weights in OpenMontage. Do not invent a Replicate, HeyGen, or Higgsfield identifier when their current public API schema does not list Seedance 2.5. +> **Resolution note.** The supported routes above expose 480p and 720p. The model +> generates 1080p natively on the vendor's own web platform, which is not an +> OpenMontage route. If a vendor blog cites 1080p or 4K, check which surface it means +> before promising it in a pipeline. + ## Reference limits - Up to 30 reference images. @@ -36,6 +41,13 @@ and `audio_urls` internally. Runway uses `references`, `referenceVideos`, and `referenceAudio`. Ark uses typed content entries with roles. Always call the OpenMontage tool instead of constructing a provider payload manually. +**50 assets is a ceiling to use deliberately, not to max out.** A cluttered reference +set with competing faces, props, and locations produces a *less* coherent result than a +smaller, chosen one. Rule: one reference per element that must stay consistent — a face, +a product, a location, a style. + +--- + ## Prompting Lead with the shot structure, then subject, action, camera, lighting, and audio. @@ -44,9 +56,206 @@ for dialogue. Reference each supplied asset by a stable role in the prompt. Thirty seconds is a ceiling, not a target: use a shorter generation when the scene has only one meaningful action. +### Section order + +One continuous block of text, broken into labeled sections. Skipping a section does not +degrade the output generally — it breaks it in a specific, predictable way. + +``` +GLOBAL STYLE + Genre, color grade, film stock or digital look, aspect ratio, shutter behavior, + and anything that must not appear. Everything below must match this. + +SCENE + A one-line logline: what happens, where, in what mood. + +CHARACTERS + Face, hair, build, wardrobe. If a reference covers this, just name which + reference is which character. + +LOCATION + The space and its props, separate from the people in it. A vague location is + the single most common reason a multi-shot sequence drifts between cuts. + +FIRST FRAME AND BLOCKING + Exact starting positions: who is where, facing which direction, as the clip + opens. Something fixed before any motion starts. + +Shot 1: [shot type] [action in one or two sentences]. Hard cut. +Shot 2: [...]. Hard cut. +Shot 3: [...]. + Pacing is built here. State camera distance and framing per shot: vague + transitions are the most common cause of a sequence collapsing. + +OPTICS / CAMERA + Focal length, camera height, handheld or dolly or crane, per shot. + +PHYSICS + Fabric, smoke, hair, liquid. Anything that would look wrong moving like a solid. + +LIGHTING + Where the light is motivated from, its direction, how it falls on faces. + +AUDIO + Ambience, specific effects, and what must not be heard. Default: + "No music, no discernible dialogue", unless the scene needs otherwise. +``` + +The shape working prompts share: **one visual rule at the top, one sound rule at the +bottom, everything in between broken into shots.** CHARACTERS can be short when a +reference does that work; keep GLOBAL STYLE, the shot breakdown, and AUDIO regardless. + +### Genre, era, and lighting carry more than prose + +On vendor web platforms these are UI selectors, and changing genre alone shifts pacing, +contrast, and camera behavior with the scene description held identical. **The API routes +above expose no such selectors**, so on OpenMontage tool calls that intent has to be +written into `GLOBAL STYLE` and `LIGHTING` explicitly — genre, decade, grain and color +response, light source and its angle, and the emotional register. Leaving them implicit is +the difference between a shot that reads as noir and one that is merely dark. + +--- + +## Locks: continuity across cuts + +Constraints do not belong in a trailing negative prompt. They go in **named blocks**, +usually phrased positively, stating what holds. + +| Lock | What it solves | +|---|---| +| `FIRST FRAME AND SPATIAL BLOCKING` | Positions as x/y percentages of frame, plus the reverse-angle note: when the camera crosses, who ends up on which side, so they never swap | +| `SCREEN DIRECTION` | The axis. "Sea always screen-right, dune line always screen-left, never reverses" | +| `LENS LOCK` | One focal per segment in degrees of diagonal field of view: 84° wide, 47° normal, 29° portrait, 18° telephoto. Plus "no lens drift mid-segment" | +| `SINGLE-LENS LOCK` | One focal for the whole piece; framing changes because the operator walks. This is what reads as documentary | +| `COUNT LOCK` | "Exactly four, never five, no background double, no mirror or reflection showing an extra copy" | +| `STILLNESS LOCK` | The list of what does not happen: nobody stands, nobody hugs, no hand leaves the blanket | +| `POSITIVE LOCKS` | Closing block restating in positive terms what holds across cuts | +| `IDENTITY / NO-IP LOCK` | "A wholly original, invented face." No logos, no readable text, no recognizable melody | + +**State monotonicity.** The cheapest continuity trick available: states only advance, +never reset. Wetness only accumulates. The cookie only shrinks — whole, bitten, two +halves, finished — and never regrows. Writing the progression shot by shot stops the model +from cleaning up your character on the next cut. + +**One grammar per action.** When an action repeats, define one way to perform it and +forbid the rest, then close the list: exactly five sword actions in the piece, each with +its timestamp. Same for slow motion — allow-list it to named moments and nowhere else. + +**Event tracks.** Write ambient elements as a timestamped event list rather than as +texture: each wave with its second and its spray height as a percentage of frame, at +irregular intervals. Works for sea, wind, rain, traffic, and crowds. + +**Substitute, don't only forbid.** If you forbid something, supply the concrete +replacement. "No blood" is weak; "the cut edge glows orange, sparks blow out, the body +crumbles to ash" is enforceable. + +--- + +## Assets + +The piece is won or lost here, before any video generates. + +**Character sheet: three plates, face on only one.** Full body front with the face +removed, full body back, and a 3/4 close-up carrying the face in two versions, smiling and +neutral. The face is stripped from the full-body plates because in a wide shot it is small +and blurry and the model copies that blur; if the only face in the set is a sharp +close-up, that is the one it uses. Without the smiling version the model invents teeth. + +**One asset per state.** Variations are split into separate assets rather than noted in +the prompt: clean jacket, soaked, and bloodied. Splitting is cheaper than arguing. + +**The rule of 10.** Before accepting an asset, generate it ten times across different poses +and lighting. It must stay recognizable all ten. + +**Locations.** Do not generate a head-on still. Generate a video of the empty space with +the camera moving slowly, then derive the other sides of the room from that footage. A +still gives you a wall; a slow travelling gives you geometry. + +--- + +## Voice + +Voice is a **conditioning sentence, not an asset**. Write it once with accent, tempo, and +manner, then paste it verbatim — without changing a word — every time that character +speaks. Rewording widens the sampling range and destabilizes the voice. + +Write accent phonetically, describing what the mouth does rather than where the speaker is +from: `th going to f and v, dropped h, glottal t, -ing to -in'`. + +Keep audio tracks separate. Never voice and music in the same reference clip. + +--- + +## What the model cannot do + +| Failure | Workaround | +|---|---| +| **Singing** | Pre-record the track, cut it into 12-second files, and label each `THE TRACK THEY ARE PERFORMING`. The label states the audio's role in the scene; `audio guide, for sync only` is technically accurate and produces nothing usable | +| **Sustained physical contact** | Fights and two bodies in constant contact. Shoot the reference on a phone and pass it as a motion reference; cheaper than twenty iterations | +| **Multi-character scenes** | The hardest category. The more faces in one shot, the more reference material each needs to avoid drifting into a generic face by the third or fourth cut | + +--- + +## Chaining and correction + +Past 30 seconds, extend the previous clip or pass it as a video reference plus a +description of what comes next — the second route also lets you introduce new characters +or objects at the seam. Avoid vendor "long video" options that stitch several generations +from one short prompt: there is not enough detail available for the runtime. + +For a single wrong element, edit rather than regenerate, stating what changes and what +holds: + +``` +Around the 2-second mark, change only [X] from [before] to [after]. +Keep the characters, movement, camera, lighting, composition and timing unchanged. +``` + +A repaired frame can also be passed as a start image for a short replacement segment. +Note the tension: a published 110-minute Seedance production banned starting frames +throughout, so continuity was carried by the assets rather than by the previous take's +last frame, which drags its own drift. Treat the start-image repair as a point fix, not as +a continuity method. + +--- + ## Cost and verification All supported routes are paid. Confirm the exact provider/model before calling and review the result for identity continuity, cuts, lip sync, audio artifacts, and prompt adherence. ComfyUI Partner Nodes use prepaid Comfy credits and are not an offline fallback. + +Iteration discipline that keeps that cost bounded: + +- **Lock the assets before generating anything.** The model has no memory: re-describe + everything, every time. +- **Change one thing per iteration**, or you will not know which change fixed it. +- **At 10 to 15 failed iterations, simplify or split the scene, not the wording.** +- Explore at the lowest resolution the route offers and only re-run what works. +- Keep a log of version, what changed, and verdict. +- **Judge the whole clip, never a single frame.** A good still can come from a generation + that falls apart at second 9. +- Land cuts on movement, never on a still pose. + +--- + +## Reference + +[`reference/techniques.md`](reference/techniques.md) — ten prompt archetypes, the settings +each used, why the combination works, and which technique each one introduces. Use it to +pick the closest archetype before writing from scratch. + +## Provenance + +The prompting method, lock patterns, and asset rules above are distilled from material the +Higgsfield team published in August 2026 — an official Seedance 2.5 prompting guide of ten +categories each tested across multiple generations, and the open-sourced production method +of a 110-minute Seedance feature. No prompt from either source is reproduced verbatim; the +technique is restated with original templates. Credit for the underlying work is theirs. + +- +- + +The route table, provider field names, and cost guidance in this file are OpenMontage's own +and take precedence over anything a vendor blog implies about API availability. diff --git a/.agents/skills/seedance-2-5/reference/techniques.md b/.agents/skills/seedance-2-5/reference/techniques.md new file mode 100644 index 00000000..33d4751f --- /dev/null +++ b/.agents/skills/seedance-2-5/reference/techniques.md @@ -0,0 +1,197 @@ +# Seedance 2.5 — the ten official prompt categories + +Higgsfield published a prompt library of ten categories in August 2026, each tested across +multiple generations to see which structures hold up on a second run. This file records the +settings each category used, why that combination works, and **which technique it +introduces**. + +Use it to pick the closest archetype before writing from scratch. + +The full prompts are at with a +copy button. They are not reproduced here — what follows is the extractable method. + +--- + +## 1. Dramatic exterior · human subject, available light + +**Settings:** Genre drama · Golden hour, single motivated source · Emotional control: grief +held beneath composure · Anamorphic large format, mixed wide and close coverage. + +**Why it works:** drama holds shots longer than most genres, so a two-hander has room to +play out. Golden hour keeps the light identical across every segment without describing the +sun's position in each one. + +**Introduces:** +- **Event track.** The sea declared "an event track, not a texture", with every wave and + gust given its timestamp and its spray height as a percentage of frame. Irregular + intervals, no two identical. +- **STILLNESS LOCK.** The list of what does not happen: nobody stands, nobody hugs, no hand + leaves the blanket. Exactly one tear falls, hers, at 13.2s. +- **Blocking in percentages.** Her at x 42%, head at y 44%. Him at x 58%, y 42%. +- **Reverse-angle note.** When the camera crosses to the seaward side she reads screen-right + and he reads screen-left. They never swap sides. +- **LENS LOCK per segment** in degrees: 47° for the wides, 29° for the close-ups. +- **Invented face.** "A wholly original, invented face", so no real likeness is cloned. + +## 2. Action sequence · multi-shot choreography + +**Settings:** Genre action · Physics realistic, high impact, real paper mass and inertia · +Anamorphic, hybrid handheld and gyro rig · Hard cold key with a single warm accent. + +**Why it works:** action tightens pacing and keeps the camera close to the movement, so a +nine-shot fight reads as one continuous sequence rather than stitched takes. + +**Introduces:** +- **Physics per material, named and marked critical.** An entire block devoted to the + banknotes being individual paper — visible fibre, worn creased edges, each bill tumbling + on its own air current — so hundreds of notes do not read as one rigid mass. +- **Marking sections "(critical)".** Flagging the block that must not be ignored. +- **Palette in percentages.** 60% cold steel, 30% matte black, 10% warm focal pop. "She and + the cash are the only warm things in a cold world." +- **Author references "in spirit, fully original execution".** Naming directors and + cinematographers without copying a specific work. +- **Explicit anti-plastic block** enumerating real optical imperfection: film grain, gate + weave, halation, chromatic aberration, lens breathing on focus pulls. +- **Cuts land on the hits**, never on a still pose. + +## 3. Commercial product · multi-reference scene + +**Settings:** Genre general · True handheld with light shake · Physics: real magic on honest +rules. + +**Why it works:** locking three references (person, location, product) keeps all three +stable while the gimmick happens around them. + +**Introduces:** the effect defined as a **mechanic with rules**, not as an adjective. Color +waves radiate from each footfall at constant radial speed, repaint what they cross, and +stay. One mechanic the model repeats identically instead of improvising it shot to shot. If +your piece has a visual trick, write it this way. + +## 4. Epic landscape · environment as protagonist + +**Settings:** Genre epic · Slow steady push-in across all three shots · Bright diffused +daylight, cold grade. + +**The shortest of the ten.** Three shots, three sentences, `Hard cut` between them, and a +closing consistent-look block with the exclusions (no people, no animals, no birds, no +shake). + +**Lesson:** with no people on screen, almost no locks are needed. Prompt complexity should +track what can actually drift, not ambition. + +## 5. Noir · shadow and atmosphere + +**Settings:** Drama with a noir base · Hard key, high contrast, monochrome · Anamorphic with +the lens locked per segment. + +**Introduces:** +- **LOCATION MAP.** The space written out: the camera works from the south sidewalk facing + north for the entire sequence, with background, midground, and foreground specified. +- **SCREEN DIRECTION LOCK.** The sedan always arrives from screen-right, the woman always + travels right-to-left, the camera never crosses to the far side of the street. +- **Lens ladder per segment:** 84° wide, 29° medium, 47° normal, 18° telephoto for the + two-shot, 84° again for the low finish, with "no lens drift mid-segment". +- **State monotonicity** applied to wetness: it only accumulates, never resets. +- **Controlled legible text:** the only readable text anywhere is the neon sign's word. + +## 6. Multi-character · consistent faces across frames + +**Settings:** Action, performance piece · Average shot under 1.5 s · Dual-quality system. + +**The hardest category:** four named characters across twenty-four fast cuts, with no time +to re-establish who is who between them. + +**Introduces:** +- **COUNT LOCK.** Exactly four, never five, no background double, no cloned member, and no + mirror, glass, or reflection showing an extra copy. +- **DUAL-QUALITY SYSTEM.** Every beat tagged `[FILM]` or `[REPORTER CAM]`, the degradation + described in full, and an explicit rule that it **never leaks** from one into the other. +- **Match-cut motif.** Round shapes and eyes that rhyme between scenes: CRT screen, + porthole, dental lamp, planet disc, pupil. +- **Wardrobe by act**, declared in the locks block. +- "Exactly five fingers per hand." + +## 7. Fantasy action · character and creature at scale + +**Settings:** Epic, action · Hyperbolic physics, no blood anywhere · Emotional control: +calm, focused, not afraid. + +**Introduces:** +- **One grammar per action.** She only cuts, never stabs, with the four phases of the cut + written out, plus a **closed list** of the five sword actions in the piece and their + timestamps. +- **Slow-motion allow-list.** Two moments only, with exact timestamps, nowhere else. +- **Substitute rather than only forbid.** No blood: the demons are cooled lava with fire + inside, the cut edge glows orange, sparks blow out, the body crumbles to ash. +- **Numeric heights.** Ledge floor is zero, the horde is 2.5 m, the archdemon 5 m with eyes + at 4.5 m; she levels at 4 m and drops to 3 m so he is looking down at her. +- **A geometric weak point** with measurements and the condition that exposes it. +- **Progressive dirt** that never cleans up. + +## 8. UGC-style ad · natural movement, no production feel + +**Settings:** Genre general, no genre-specific visual logic imposed · Handheld selfie framing +fixed for the whole video · Soft even cabin light · Emotional control: understated, no +performing. + +**Introduces:** +- **Performance by imperfection.** Natural blink rate, reactions arriving half a beat late, + smiles that start small and grow, the glance away from the lens and back mid-sentence, the + lowered voice of someone self-conscious about filming on a plane. +- **Live background people, not NPCs.** Each with an independent activity; nobody looks at + the camera and nobody reacts in sync. +- **Fixed seating geometry** across all eight shots. +- **State monotonicity** applied to the cookie: whole, bitten, two halves, finished, never + regrowing. +- **Turbulence that exists in two shots only and never returns.** +- **Dialogue per shot in complete sentences**, with all speech ending by 25s. + +## 9. Horror · wrong light, wrong angle + +**Settings:** Genre horror · Deep one-point perspective, 4:3 · Real weight and inertia · +Emotional control: tired, welling tears, held tension without release. + +**Introduces:** the negative space past the shoulder declared first empty, then filled — +dread built from absence rather than a reveal. + +And the most elegant rule of the ten: the antagonist is **never in sharp focus** until the +final second. Always a blur, a silhouette, a fragment behind shelves. "The autofocus refuses +to lock on her." Focus directed as a narrative device. + +## 10. Documentary · observational camera + +**Settings:** General, third-person observational · A single 47° lens for the entire +sequence · Five distinct natural light states across one continuous span of time. + +**Introduces:** +- **SINGLE-LENS LOCK.** One focal length for thirty seconds; every framing change is + achieved by the operator physically walking, with camera-to-subject distance declared per + segment: 5 m, 2 m, 3 m, 2.5 m, 7 m. +- **Crowd size as the clock.** Forty people at the start, sixty, a hundred at the peak, + fifteen or twenty deep in the night, five or six by dawn — so it reads as one continuous + night rather than five separate scenes. +- **FRAMING RESPECT LOCK.** The camera treats her as a person at a party, not a body: no + slow pans up or down a figure, no isolated shots of torso or legs, no low angles. +- **BEAUTY WITHOUT RETOUCHING LOCK.** Visible pores, developing sunburn, salt drying to a + faint bloom, sand along the forearms, tired eyes by dawn. +- **CLEAN-AIR LOCK.** Nobody smokes; the only smoke comes from the fires and always travels + out to sea. +- **Wardrobe layered progressively** as the temperature drops. + +--- + +## The six findings from their testing + +1. **References hold identity across cuts more reliably than a text description.** Once a + face or product is locked as a reference it survives hard cuts, lighting changes, and + camera moves that would otherwise cause drift. +2. **Genre and lighting settings do more work than most of the prompt text.** +3. **Multi-character scenes are the hardest category to keep consistent.** The more faces in + one shot, the more reference material each needs. +4. **The shot-by-shot breakdown outperforms a single continuous scene description** for + anything with more than one camera angle. Vague transitions are the most common cause of + a sequence falling apart. +5. **50 references is a ceiling worth using deliberately, not maxing out by default.** A + cluttered set produces a less coherent result than a smaller, carefully chosen one. +6. **GLOBAL STYLE and AUDIO function as guardrails more than instructions.** Explicitly + excluding what should not appear prevents more failures than adding positive description. diff --git a/skills/INDEX.md b/skills/INDEX.md index a5447afd..68959d0f 100644 --- a/skills/INDEX.md +++ b/skills/INDEX.md @@ -325,5 +325,5 @@ Claude Code accesses them via symlinks in `.claude/skills/`. | **Design** | `tailwind-design-system`, `web-design-guidelines`, `vercel-react-best-practices`, `vercel-composition-patterns` | `wshobson/agents`, `vercel-labs/agent-skills` | | **AI Video (HeyGen)** | `heygen`, `avatar-video`, `create-video`, `faceswap`, `ai-video-gen`, `video-download`, `video-edit`, `video-translate`, `video-understand`, `visual-style` | `heygen-com/skills` | | **AI Video/Image/TTS/Avatar (Kling Official)** | `kling-official` - official direct API auth, Classic/Turbo/Omni task protocols, multi-reference Omni syntax, internal Elements/Account Usage helpers, callback notes, TTS voice parameters, avatar/lip-sync face selection, error handling, and cost governance for `kling_official_video` / `kling_official_image` / `kling_tts` / `kling_avatar` / `kling_lip_sync` | Local OpenMontage skill | -| **AI Video (Premium)** | `seedance-2-0` — preferred premium default (cinematic, trailer, multi-shot, lip-sync, synced audio); accessed via `seedance_video` (fal.ai) or `heygen_video` Avatar Shots | Local OpenMontage skill | +| **AI Video (Premium)** | `seedance-2-0` — preferred premium default (cinematic, trailer, multi-shot, lip-sync, synced audio); accessed via `seedance_video` (fal.ai) or `heygen_video` Avatar Shots; `seedance-2-5` — 4–30 s clips, 50 multimodal references, prompt contract (section order, `Hard cut` breakdown, continuity locks, asset method) | Local OpenMontage skill | | **Infrastructure** | `acestep`, `ltx2`, `playwright-recording` | `digitalsamba/claude-code-video-toolkit` |