Commit Graph

100 Commits

Author SHA1 Message Date
calesthio
00e4902348 Merge remote-tracking branch 'origin/main' into codex/repair-pr-371
# Conflicts:
#	docs/PROVIDERS.md
2026-08-13 11:20:19 -07:00
calesthio
329628a22b Merge remote-tracking branch 'origin/main' into codex/repair-pr-371 2026-08-13 11:11:02 -07:00
calesthio
e2a9e35c40 Merge remote-tracking branch 'origin/main' into codex/repair-pr-442 2026-08-13 11:11:02 -07:00
calesthio
6b722ab56e Merge remote-tracking branch 'origin/main' into codex/repair-pr-458
# Conflicts:
#	tools/graphics/image_selector.py
2026-08-13 10:38:22 -07:00
Calesthio
0adcc23f44 Merge pull request #457 from wulongshe/feat/gemini-flash-image
feat: add gemini-2.5-flash-image backend to google_imagen
2026-08-13 10:36:12 -07:00
calesthio
a5c9d789be Merge remote-tracking branch 'origin/main' into codex/repair-pr-458
# Conflicts:
#	tools/graphics/image_selector.py
2026-08-13 10:17:12 -07:00
calesthio
700046d33c Merge remote-tracking branch 'origin/main' into codex/repair-pr-482
# Conflicts:
#	docs/PROVIDERS.md
2026-08-13 10:17:12 -07:00
calesthio
4e5bd8d5fa Merge remote-tracking branch 'origin/main' into codex/repair-pr-457
# Conflicts:
#	tools/graphics/image_selector.py
2026-08-13 10:17:11 -07:00
calesthio
be8cf96861 Merge remote-tracking branch 'origin/main' into codex/repair-pr-371 2026-08-13 10:15:43 -07:00
calesthio
0bb4c266d7 Merge remote-tracking branch 'origin/main' into codex/repair-pr-442 2026-08-13 10:15:35 -07:00
Calesthio
5fcea90d71 Merge pull request #483 from ikohu-66/add_seedream_tools
Add seedream tools
2026-08-13 10:08:25 -07:00
calesthio
f21e86570d Merge remote-tracking branch 'origin/main' into codex/repair-pr-458 2026-08-13 09:44:58 -07:00
calesthio
6b75448a0f Merge remote-tracking branch 'origin/main' into codex/repair-pr-483 2026-08-13 09:44:57 -07:00
calesthio
3f2898c77e Merge remote-tracking branch 'origin/main' into codex/repair-pr-495 2026-08-13 09:44:57 -07:00
calesthio
e614c9d4c4 Merge remote-tracking branch 'origin/main' into codex/repair-pr-457 2026-08-13 09:44:56 -07:00
calesthio
078eb7eecd fix: integrate Azure TTS with shared selector 2026-08-13 09:27:32 -07:00
calesthio
a3c45aa3bf Merge remote-tracking branch 'origin/main' into codex/repair-pr-371
# Conflicts:
#	.env.example
2026-08-13 09:23:46 -07:00
calesthio
cffc18308f fix: scope fal providers and Google TTS networking 2026-08-13 09:22:11 -07:00
calesthio
b3434affe3 Merge remote-tracking branch 'origin/main' into codex/repair-pr-482
# Conflicts:
#	skills/pipelines/animation/asset-director.md
2026-08-13 09:20:01 -07:00
calesthio
8143266ead fix: complete Hunyuan image provider integration 2026-08-13 09:18:49 -07:00
calesthio
171866dbbf fix: complete Seedream provider integration 2026-08-13 09:12:00 -07:00
calesthio
bf55d498fb Merge remote-tracking branch 'origin/main' into codex/repair-pr-442
# Conflicts:
#	.env.example
2026-08-13 09:07:59 -07:00
calesthio
b1aefb364a fix: complete MiniMax image provider contracts 2026-08-13 09:02:22 -07:00
calesthio
579bf053e7 fix: route Gemini image models through selector 2026-08-13 08:58:27 -07:00
calesthio
04571cfe8e test: make Blender doctor contract portable 2026-08-13 07:56:45 -07:00
calesthio
2702362c24 feat: add production 3D world pipeline 2026-08-13 07:48:38 -07:00
octo-patch
9c2850f02e feat: add MiniMax image generation tool 2026-08-12 22:46:06 +08:00
蓝友和
63fd646717 add seedreamm tools 2026-08-09 00:25:44 +08:00
Codex
21d51ab9c8 add fal ElevenLabs speech and secure audio routing 2026-08-08 13:02:07 +00:00
Codex
6b4c73c6df scope Google TTS credentials safely 2026-08-07 15:01:59 +00:00
calesthio
9482eddeff fix: recover bounded defects from PR backlog 2026-08-03 02:14:01 -07:00
clarkh
00745a5f6d feat: add Hunyuan Image Generation 3.0 (混元生图) via TokenHub API
Async text-to-image tool using Tencent TokenHub (hy-image-v3.0).
Parameters mirror upstream SubmitTextToImageJob API: prompt, resolution,
seed, revise, logo_add, logo_param, and reference images via Images.N.

- tools/graphics/hunyuan_image.py: submit → poll → download flow,
  matching hunyuan_cloud_video.py code style
- tests/tools/test_hunyuan_image.py: 32 unit tests covering payload
  building, image resolution, API error handling, and mocked e2e flow
2026-07-30 11:38:33 +08:00
shewulong
5d152e4699 feat: add gemini-2.5-flash-image backend to google_imagen 2026-07-29 20:52:05 -05:00
zzzxtnt
b71f9b1f3e feat(video): add direct Volcengine Ark Seedance 2.0 provider 2026-07-27 11:24:07 +08:00
Calesthio
888fe5a73c Merge pull request #337 from 0xDevNinja/fix/green-screen-chromakey-1x1-collapse
fix(green_screen): scale chromakey background to frame size, not 1x1
2026-07-21 14:57:37 -07:00
Calesthio
33c7000e68 Merge pull request #374 from ShiroKSH/fix/hyperframes-relative-output-path
fix: resolve relative HyperFrames output paths
2026-07-17 18:53:52 -07:00
Calesthio
2a6bf1e039 Merge pull request #391 from tianrking/agent/fix-delayed-audio-fades
fix(audio): schedule delayed track fades correctly
2026-07-17 18:51:56 -07:00
tianrking
5e21f1b78a fix(audio): schedule delayed track fades correctly 2026-07-16 13:33:17 +08:00
0xDevNinja
011a27df6a fix(google_imagen): write every generated image, not just the first
execute() sends sampleCount=number_of_images and estimate_cost() bills
0.04 * n, but result handling decoded only predictions[0] and wrote it to
a single output_path. Images 2..n were dropped: never decoded, never
written, absent from artifacts. The user paid for n and received one.

The result also misreported the drop rather than failing loudly --
images_generated returned len(predictions) (what the API sent) while
artifacts held a single path, so an agent picking between variants read a
count that did not match the artifact list.

Add _output_paths() and loop over every prediction, mirroring the pattern
already used by openai_image and grok_image: suffix multi-image paths
_1/_2/... so none overwrite each other, keep the exact requested path when
n=1, return all paths in artifacts, and report images_generated as the
count actually written.

Closes #388
2026-07-15 16:30:57 +05:30
0xDevNinja
fa756fbec5 fix(green_screen): make chromakey compositing portable across FFmpeg builds
The CI Linux FFmpeg build carried the keyed frame forward without an alpha
plane, so overlay drew opaque green over the background (corner stayed green)
instead of compositing — the E2E test failed there even though it passed on
macOS/Windows.

Force `format=yuva420p` immediately after chromakey so the keyed transparency
always has an explicit alpha plane, and size the background to the frame up
front (color=...:size=WxH, passing the probed width/height into
_process_chromakey) instead of scaling a 1x1 source with scale2ref — dropping
scale2ref also removes the format negotiation that discarded the alpha on some
builds. Output is flattened to yuv420p after the overlay.
2026-07-14 14:27:40 +05:30
ShiroKSH
24617af460 fix: resolve relative HyperFrames output paths 2026-07-13 21:33:11 +03:00
amartya-dev
888d7b1e72 feat(tts): add Azure AI Speech as an optional cloud text-to-speech provider
Neural TTS via the synchronous REST v1 endpoint (SSML body, no token
exchange or job polling). Shares one Speech resource with azure_stt —
AZURE_SPEECH_KEY + AZURE_SPEECH_REGION unlock both directions; optional
AZURE_TTS_ENDPOINT overrides the TTS host (a different subdomain than
the STT endpoint). piper_tts remains the default offline path.

- tools/audio/azure_tts.py: azure_tts tool (capability=tts), voice
  shortlist aliases, SSML prosody/style, mp3/wav output, cost tracking
- tests/tools/test_azure_tts.py: contract, discovery, status, SSML,
  and mocked execute tests (21 tests, no live network)
- .agents/.claude skills: azure-text-to-speech Layer-3 skill
- docs: PROVIDERS.md section + tables, ARCHITECTURE.md inventories,
  AGENT_GUIDE.md + skills/INDEX.md rows, asset-director TTS cheatsheet,
  .env.example
2026-07-13 20:04:28 +05:30
Calesthio
f8d94632ea Merge pull request #354 from amartya-dev/feat/azure-speech-to-text
feat(stt): add Azure AI Speech as an optional cloud speech-to-text provider
2026-07-12 10:48:36 -07:00
Calesthio
318e0843ee Merge pull request #325 from 0xDevNinja/fix/subtitle-ts-overflow-and-checkpoint-keyerror
fix: subtitle timestamp ms overflow; checkpoint KeyError on manifest-only stages
2026-07-12 10:45:47 -07:00
amartya-dev
a2a0d8c8af feat(stt): add Azure AI Speech as an optional cloud speech-to-text provider
Add an Azure AI Speech transcription tool. It is opt-in: when
AZURE_SPEECH_KEY is configured the agent may prefer it for cloud STT,
while the local faster-whisper `transcriber` stays the default offline
path. Shared pipeline manifests are intentionally left unchanged, so no
default provider selection is altered for existing users.

- tools/analysis/azure_stt.py: new `azure_stt` tool (capability=analysis,
  provider=azure) calling the Fast Transcription REST API. The local file
  is uploaded via multipart and transcribed synchronously with word-level
  timestamps and optional diarization — no Blob storage or async polling.
  Output schema mirrors `transcriber` exactly, so it is a drop-in for
  `subtitle_gen` and other transcript consumers. Follows the existing
  provider-tool conventions (env-var status check, `_transcribe` helper,
  cost_usd/model on the result, fallback="transcriber").
- Auto-discovered by the registry; no registry or selector changes.
- tests/tools/test_azure_stt.py: contract, discovery, status, response
  mapping, execute guardrails, and a mocked-network success path (no live
  API calls).
- .agents/skills + .claude/skills: azure-speech-to-text Layer-3 skill.
- docs/PROVIDERS.md: Azure AI Speech setup, API notes, and pricing.
- .env.example, skills/INDEX.md, AGENT_GUIDE.md: document the optional
  cloud STT path alongside the default whisper transcriber.
2026-07-10 23:30:19 +05:30
0xDevNinja
df8cf21310 fix(green_screen): scale chromakey background to frame size, not 1x1
_process_chromakey built the composite background from a 1x1 lavfi color source
and tried to size it with `[0:v]scale=iw:ih`. That scale is a no-op — iw/ih are
the 1x1 source's own dimensions, and there is no cross-reference to the frame.
FFmpeg's overlay then takes the size of its first input (the 1x1 background), so
every processed frame is clipped to a single pixel. The exception fallback never
runs because the primary command exits 0 (a valid 1x1 PNG), and
_reconstruct_video upscales those 1x1 frames — producing a solid-color video
with the keyed subject entirely gone. Total data loss for method="chromakey"
(and method="auto" when it selects chromakey).

Use scale2ref to resize the background to the actual frame dimensions before
overlaying, so the keyed subject is composited at full resolution.

Verified with ffmpeg: a 320x240 green frame with a red subject now produces a
320x240 output with the subject preserved and green replaced by the background,
instead of a 1x1 (then upscaled solid-color) frame.
2026-07-09 18:15:05 +05:30
calesthio
2ef18e77a9 fix(video): normalize Gemini Omni file URIs; document provider in PROVIDERS.md
Review findings from PR #333:

P1: _download_via_uri assumed output_video.uri is always files/<id>.
The API can return a full resource URI or a ready-made
.../files/<id>:download?alt=media download URL, which produced an
invalid poll path with a second :download appended. New
_file_id_from_uri() extracts the bare id from every documented shape;
regression tests cover the full-URL form plus a parametrized matrix of
URI shapes.

P2: docs/PROVIDERS.md still described the Google key as TTS + Imagen
only. The shared-key section now covers gemini_omni_video (model id,
~$0.10/sec pricing table, paid-tier-only, edit-turn billing note), and
the env snippet, provider-to-tool mapping, and capability coverage
tables include the new provider.
2026-07-08 23:45:47 -07:00
calesthio
34d1053526 feat(video): add Gemini Omni Flash provider with conversational editing
Add gemini_omni_video, a native Gemini API provider wrapping
gemini-omni-flash-preview via the Interactions API. Text-to-video,
image/reference-to-video with <FIRST_FRAME>/<IMAGE_REF_N> prompt tags,
and stateful edit_video turns via previous_interaction_id — the only
provider in the fleet that can refine a clip without regenerating it.
Reuses the existing GOOGLE_API_KEY / GEMINI_API_KEY, so one Google key
now unlocks images, TTS, and video.

- New Layer 3 skill .agents/skills/gemini-omni (prompting, edit-loop
  rules, tag/timecode syntax, preview limits) sourced from official
  Google docs; linked via agent_skills and the AGENT_GUIDE Layer 3 map
- ai-video-gen gains the Gemini API gateway row + editing pointer
- veo_video/sora_video fallback lists and video_selector agent_skills
  reference the new provider; quality_score 0.85 with rationale
- Contract tests: registry discovery, selector routing, status from
  env keys, uri + inline delivery, edit turns, typed image parts,
  store=false editability, cost clamp
2026-07-08 11:00:57 -07:00
0xDevNinja
bcd8eb6e53 fix(subtitle_gen): stop millisecond rounding from overflowing timestamps
_ts_srt/_ts_vtt computed the seconds and millisecond fields independently:
`ms = int(round((seconds % 1) * 1000))`. When the fractional part is >= 0.9995
that rounds to 1000, emitting a malformed 4-digit `…,1000` value with no carry
into the seconds field (and, at 59.9999/3599.9999, no carry into minutes/hours).
For example 0.9999s became `00:00:00,1000` instead of `00:00:01,000`. ASR word
and segment end-times routinely land on such fractional boundaries, and the
resulting cue is rejected or mistimed by strict SRT/VTT parsers (ffmpeg
subtitles filter, VLC, browser WebVTT).

Decompose from a single rounded total-milliseconds value so the carry
propagates across all fields. Both formatters now share one `_hmsms` helper.
2026-07-07 15:29:03 +05:30
calesthio
015809c103 fix(ci): isolate requests module tests 2026-07-06 22:36:03 -07:00