Commit Graph

402 Commits

Author SHA1 Message Date
calesthio
bf55d498fb Merge remote-tracking branch 'origin/main' into codex/repair-pr-442
# Conflicts:
#	.env.example
2026-08-13 09:07:59 -07:00
calesthio
b1aefb364a fix: complete MiniMax image provider contracts 2026-08-13 09:02:22 -07:00
calesthio
579bf053e7 fix: route Gemini image models through selector 2026-08-13 08:58:27 -07:00
Calesthio
51330d1723 Merge pull request #499 from calesthio/codex/paper-3d-world
feat: add production 3D world pipeline
2026-08-13 08:01:20 -07:00
calesthio
04571cfe8e test: make Blender doctor contract portable 2026-08-13 07:56:45 -07:00
calesthio
2702362c24 feat: add production 3D world pipeline 2026-08-13 07:48:38 -07:00
octo-patch
9c2850f02e feat: add MiniMax image generation tool 2026-08-12 22:46:06 +08:00
clarkh
b6dded18fe Merge branch 'main' into feat/hunyuan_cloud_video 2026-08-12 21:24:23 +08:00
蓝友和
63fd646717 add seedreamm tools 2026-08-09 00:25:44 +08:00
Codex
2ed75c5a25 document Google TTS IPv4 restriction switch 2026-08-08 13:12:19 +00:00
Codex
21d51ab9c8 add fal ElevenLabs speech and secure audio routing 2026-08-08 13:02:07 +00:00
Codex
6b4c73c6df scope Google TTS credentials safely 2026-08-07 15:01:59 +00:00
Ntsako
cf1a13a722 feat: bundle a native-node ACE-Step v1 workflow for comfyui_music
ACE-Step v1's node-pack fragmentation turns out to be moot: ComfyUI ships
TextEncodeAceStepAudio/EmptyAceStepLatentAudio as native core nodes
(comfy_extras/nodes_ace.py), not a third-party pack, and Comfy-Org's own
workflow_templates repo has an official ACE-Step-v1 template built from
those plus long-stable core nodes. tools/_comfyui/workflows/ace-step-1-t2a.json
was built by cross-checking every node's class_type and input names against
ComfyUI's own source (nodes_ace.py, nodes_audio.py, nodes_latent.py,
nodes.py) rather than trusting the UI-format export directly.

comfyui_music now defaults to this bundled workflow: prompt maps to
ACE-Step's tags field (matching suno_music's "prompt = music description"
convention), lyrics/duration_seconds/steps/cfg/lyrics_strength/seed are all
patchable, and missing ace_step_v1_3.5b.safetensors surfaces through the
same missing_models contract as image/video. workflow_json/workflow_path +
output_node remains available for ACE-Step 1.5, other node packs, or
different audio models entirely.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-06 15:10:48 +02:00
Ntsako
172acca6ea feat: ship comfyui_music as a custom-workflow-only ACE-Step tool
Resolves the "music generation" open question from the adapter plan.
Unlike comfyui_image/comfyui_video there is no bundled workflow: ACE-Step's
ComfyUI node interface isn't standardized across custom node packs
(AceStepModelLoader vs native TextEncodeAceStepAudio, etc.), so instead of
picking one pack and breaking for everyone else, comfyui_music always
requires a caller-supplied workflow_json/workflow_path + output_node --
the same override contract image/video offer as an alternative, just
mandatory here. prompt is provenance-only, never injected into the graph.

Routed through the existing registry.get_by_capability("music_generation")
path alongside suno_music/music_gen -- no dedicated selector needed.
ComfyUIClient.generate() now also reads the "audio" output key (what
ComfyUI's native SaveAudio node writes), and gets timeout/resume/websocket-
wait/multi-server support for free via the shared client. Duration is a
best-effort ffprobe probe of the downloaded file since a custom workflow
gives no other way to know it ahead of time.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-06 14:19:05 +02:00
Ntsako
2f114682e8 feat: support per-capability ComfyUI server URLs for image/video
Resolves the "multi-server" open question from the adapter plan.
ComfyUIClient(capability="image"|"video") now resolves its server URL
from COMFYUI_IMAGE_SERVER_URL / COMFYUI_VIDEO_SERVER_URL first, falling
back to the shared COMFYUI_SERVER_URL and then the localhost default —
so comfyui_image and comfyui_video can point at separate ComfyUI
instances (different GPUs, different model sets) with zero extra config
for single-server setups. is_default_url/unavailable_reason() and the
setup_offer metadata account for the override.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-06 12:49:57 +02:00
Ntsako
ca203e49b7 feat: wait on ComfyUI websocket feed instead of polling for completion
Resolves the "async generation" open question from the adapter plan.
generate() now watches ComfyUI's websocket events (executing/progress/
execution_error) and reacts immediately instead of sleeping between REST
polls, with an optional on_progress callback that comfyui_video uses to
print step progress on long renders. websocket-client is an optional
import; _wait() falls back to the original poll() loop (with the
remaining time budget, not a fresh one) when it's unavailable or the
connection drops, so resume_prompt_id recovery is unaffected either way.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-06 12:31:24 +02:00
Ntsako
ceee7c7d56 fix: raise ComfyUI video timeout default and add job resume support
Non-accelerated local-GPU workflows (e.g. Wan 1.3B at 832x480/81-97
frames) routinely took ~1360-1630s, so the old 900s default false-failed
real renders that were still completing server-side. Timeout is now a
configurable timeout_seconds input (default 3600s), and ComfyUIError
carries the prompt_id on error/timeout so a timed-out-but-still-running
job can be resumed via resume_prompt_id instead of resubmitted.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-06 10:55:56 +02:00
Calesthio
4eab34c5cf Merge pull request #466 from calesthio/codex/backlog-recovery-20260802
Recover bounded defects from the latest PR backlog
2026-08-03 14:49:08 +05:30
calesthio
9482eddeff fix: recover bounded defects from PR backlog 2026-08-03 02:14:01 -07:00
蓝友和
cf3407155c add seedream v5 tools for fai.ai 2026-08-03 00:53:20 +08:00
蓝友和
efb86f55dd add seedream tool 2026-08-02 18:23:53 +08:00
clarkh
00745a5f6d feat: add Hunyuan Image Generation 3.0 (混元生图) via TokenHub API
Async text-to-image tool using Tencent TokenHub (hy-image-v3.0).
Parameters mirror upstream SubmitTextToImageJob API: prompt, resolution,
seed, revise, logo_add, logo_param, and reference images via Images.N.

- tools/graphics/hunyuan_image.py: submit → poll → download flow,
  matching hunyuan_cloud_video.py code style
- tests/tools/test_hunyuan_image.py: 32 unit tests covering payload
  building, image resolution, API error handling, and mocked e2e flow
2026-07-30 11:38:33 +08:00
clarkh
2237a9bcf9 refactor: remove TENCENT_TOKENHUB_MODEL env var, relocate Hunyuan entry in .env.example
The env var added complexity without meaningful benefit — explicit model
selection via the "model" input parameter is sufficient for both image and
video TokenHub tools.

- tools/video/hunyuan_cloud_video.py: drop env var fallback from
  _resolve_model(), remove mention from install_instructions
- tests/contracts/test_hunyuan_cloud_video.py: remove
  test_resolve_model_from_env and test_resolve_model_input_overrides_env
- .env.example: move TENCENT_TOKENHUB_API_KEY to a dedicated "Tencent
  Hunyuan TokenHub API" section, drop TENCENT_TOKENHUB_MODEL comment,
  broaden description from "video generation" to generic "Tencent Hunyuan
  via TokenHub API"
2026-07-30 10:53:33 +08:00
shewulong
5d152e4699 feat: add gemini-2.5-flash-image backend to google_imagen 2026-07-29 20:52:05 -05:00
clarkh
287c77fa64 feat: add Tencent Hunyuan cloud video provider via TokenHub API
Introduce a new video generation provider backed by the Tencent TokenHub
API (tokenhub.tencentmaas.com), an OpenAI-compatible gateway for Tencent
Hunyuan video models with simple Bearer-token auth.

- Add hunyuan_cloud_video tool (submit → poll → download) supporting
  both text-to-video (hy-video-1.5) and image-to-video (yt-video-2.0)
- Add env vars: TENCENT_TOKENHUB_API_KEY, TENCENT_TOKENHUB_MODEL
- Add contract tests for the new tool
- Document setup, API flow, model pricing, and schema constraints in
  PROVIDERS.md
- Update provider tables and capability matrix throughout docs
2026-07-28 16:48:06 +08:00
zzzxtnt
b71f9b1f3e feat(video): add direct Volcengine Ark Seedance 2.0 provider 2026-07-27 11:24:07 +08:00
calesthio
c36e41223e Make Monty the Clapper the OpenMontage mascot
Replace the play-button logo in both READMEs with animated SVG versions of
Monty, served via <picture> + prefers-color-scheme so the mark reads on
GitHub's light and dark themes. Motion uses SMIL animateTransform rather
than CSS keyframes so it survives the <img> rendering context without
depending on transform-box: view-box.

Rebuild the 1280x640 social preview around Monty in the ink/cream/terracotta
palette, and refresh its stat row to the current counts (12 pipelines,
100+ tools, 700+ agent skills). The card's source HTML ships alongside it so
future count updates are an edit and a re-screenshot.
2026-07-24 12:25:41 -07:00
Calesthio
0af32ce5e1 Merge pull request #341 from yiyabo/feat/jimeng-video
Add Volcengine Jimeng (即梦 AI) video provider with V4 signing
2026-07-23 07:59:05 -07:00
Yiyabo
d4426f6e94 fix: declare env dependencies and map selector duration to frames
- Add env:VOLC_ACCESSKEY and env:VOLC_SECRETKEY to dependencies
- Map selector 'duration' (seconds) to Jimeng 'frames' (121/241)
- Add 4 selector duration mapping regression tests
- 59 tests pass

Fixes calesthio's third review feedback on PR #341.
2026-07-23 01:33:36 +08:00
calesthio
e0bfcb44a6 Update README capability counts 2026-07-22 07:16:53 -07:00
Tomofumi Yagi
990d7f9a2c fix: address selector contract, promo pricing, idempotency, and registry metadata review
- Require 'model' in the input schema and accept 'model_id' as a
  selector-compatible alias (tts_selector exposes model_id); add a
  selector-routing regression test
- Document s2.1-pro-free as promotional (free through end of July 2026,
  Fair Use, no SLA, possible request retention, commercial-use
  restrictions) in PROVIDERS.md and the Layer 3 skill; estimate_cost()
  falls back to the paid s2.1-pro rate after the promo window
- Normalize voice_id/reference_id and model_id/model aliases before
  computing the idempotency key, and include all output-affecting inputs
  (bitrate, sample_rate, temperature, top_p, repetition_penalty, latency,
  prosody, normalize, chunk_length) with API defaults applied
- Declare env:FISH_AUDIO_API_KEY in dependencies so registry metadata
  reports the requirement
- Add fish_audio to the TTS provider set in the phase3 registry contract
  test
2026-07-22 19:00:46 +09:00
Tomofumi Yagi
b29e238e8e docs: add fish.audio to provider docs and fix stale speech-1.x reference
- .env.example: replace removed speech-1.x mention with the actual
  supported backends (s1 / s2-pro / s2.1-pro)
- docs/PROVIDERS.md: add fish.audio section (setup, backend models,
  per-byte pricing incl. free s2.1-pro-free tier) plus entries in the
  env var summary, Provider-to-Tool Mapping, and Capability Coverage
- skills/INDEX.md: list fish-audio-tts in the TTS & Audio Layer 3 row
2026-07-22 18:52:33 +09:00
Tomofumi Yagi
d40c32441c feat: add fish.audio TTS provider
Add FishAudioTTS (capability=tts) so tts_selector auto-discovers a new
high-quality, voice-clone-capable provider. Backend model is required per
call: s1 (previous flagship, kept for compatibility), s2-pro (first S2
generation), s2.1-pro (latest flagship — inline emotion tags, 80+
languages), s2.1-pro-free (free tier for drafts). s1-mini and the
speech-1.x tier have been removed from the current fish.audio API and are
no longer supported. Voice cloning via reference_id with voice_id as a
selector-compatible alias. Adds temperature/top_p/repetition_penalty
sampling controls, optional sample_rate, opus output format, and a "low"
latency tier. Cost is estimated per UTF-8 byte to match fish.audio
billing. Includes a Layer 3 skill, .env.example entry, and unit tests.

Verified end-to-end with s2.1-pro + reference_id: generated a 7-segment
Japanese narration successfully.
2026-07-22 18:51:05 +09:00
Calesthio
888fe5a73c Merge pull request #337 from 0xDevNinja/fix/green-screen-chromakey-1x1-collapse
fix(green_screen): scale chromakey background to frame size, not 1x1
2026-07-21 14:57:37 -07:00
Calesthio
c934116c0a Merge pull request #413 from albatrossflyon-coder/fix/skill-template-os-command-injection
fix: sanitize os.system() shell injection in manimgl scene templates
2026-07-21 14:55:43 -07:00
Chris Brown
85a63471a2 fix: sanitize os.system() shell injection in manimgl scene templates
os.system(f"manimgl {__file__} ClassName") interpolates the script's own
path into a shell string. These templates are meant to be copied and
renamed per-scene by an agent, so a scene/folder name containing shell
metacharacters is a real injection path, not just malformed input.

Switched to subprocess.run() with an argument list (no shell=True), so
there's nothing left for a shell to interpret regardless of what the
path contains. Same fix applied in both duplicate locations
(.claude/skills and .agents/skills) since the files are identical.

Reviewed scripts/lib/tts.mjs's child_process usage as part of the same
report -- not included in this PR, it already passes args as a real
array with no shell:true anywhere in the call chain, so it isn't
actually exploitable.
2026-07-21 11:26:42 -05:00
calesthio
db91727598 Add website link to README
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-19 15:06:44 -07:00
Calesthio
80e045a1e7 Merge pull request #404 from calesthio/codex/remove-cla
Remove the CLA contribution gate
2026-07-19 09:24:36 -07:00
calesthio
8190911ca8 chore: remove CLA contribution gate 2026-07-19 09:16:31 -07:00
Yiyabo
53773cd5fc fix: tighten input_schema per Volcengine Jimeng 3.0 Pro contract
- frames: enum [121, 241] (was minimum 1)
- prompt: maxLength 800 (was 2000)
- seed: minimum -1 (was unbounded)
- Add 9 schema validation rejection tests
- Add authoritative API reference link to PROVIDERS.md
- Document CVSync2Async* route choice and schema constraints

Fixes calesthio's second review feedback on PR #341.
2026-07-19 21:47:06 +08:00
calesthio
c2045ad5f0 Merge Lyria music generation skill 2026-07-18 17:01:58 -07:00
Calesthio
6493ca6ed3 Merge pull request #401 from calesthio/cla-setup
Add Contributor License Agreement and CLA signature gate
2026-07-18 16:48:15 -07:00
calesthio
0ff7bb30e3 feat(music): add Google Lyria generation skill 2026-07-18 16:47:35 -07:00
calesthio
9f8f99e97e Add Contributor License Agreement and CLA signature gate
Introduces CLA.md (individual CLA: contributors keep all rights to their
work, grant the project the right to offer contributions under additional
license terms; includes a written commitment in section 6 that the engine
remains open source) and a CLA Assistant Lite workflow that gates every PR
on a one-comment signature, with signatures stored in-repo on the
cla-signatures branch.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-18 16:44:11 -07:00
calesthio
af87fc1337 fix(backlot): surface approval artifacts before gates 2026-07-18 01:33:12 -07:00
calesthio
5072b4647c docs: update Bloome sponsor links 2026-07-17 23:10:34 -07:00
Calesthio
33c7000e68 Merge pull request #374 from ShiroKSH/fix/hyperframes-relative-output-path
fix: resolve relative HyperFrames output paths
2026-07-17 18:53:52 -07:00
Calesthio
2a6bf1e039 Merge pull request #391 from tianrking/agent/fix-delayed-audio-fades
fix(audio): schedule delayed track fades correctly
2026-07-17 18:51:56 -07:00
Calesthio
9f94c1fe0e Merge pull request #389 from 0xDevNinja/fix/imagen-multi-image-drop
fix(google_imagen): write every generated image, not just the first
2026-07-17 18:50:08 -07:00
Calesthio
39e46d78fe Merge pull request #393 from 0xDevNinja/fix/corpus-diversify-scale
fix(corpus): normalize diversify() position term so similarity can compete
2026-07-17 18:48:13 -07:00