Go: implement TTS, ASR for Siliconflow and TTs for StepFun (#14944)

### What problem does this PR solve?

This PRimplement TTS, ASR for Siliconflow and TTs for StepFun

**The following functionalities are now supported:**

**SiliConFlow:**
- [x] Text To Speech
- [x] Audio To Text
- [x] Stream Audio To Text

**StrepFun:**

- [x] Audio To Text
- [x] Stream Audio To Text

**Verified examples from the CLI:**
```plaintext
# SiliconFlow

RAGFlow(user)> tts with 'FunAudioLLM/CosyVoice2-0.5B@test@Siliconflow' text 'hello? show yourself' play format 'wav' param '{"voice": "fnlp/MOSS-TTSD-v0.5:alex"}'
SUCCESS

RAGFlow(user)> asr with 'FunAudioLLM/SenseVoiceSmall@test@siliconflow' audio './internal/test.wav' param ''
+----------------------------------------------------------------------------------------------------------------------+
| text                                                                                                                 |
+----------------------------------------------------------------------------------------------------------------------+
| The examination and testimony of the experts enabled the commission to conclude that five shots may have been fired. |
+----------------------------------------------------------------------------------------------------------------------+

RAGFlow(user)> stream asr with 'FunAudioLLM/SenseVoiceSmall@test@siliconflow' audio './internal/test.wav' param ''
+----------------------------------------------------------------------------------------------------------------------+
| text                                                                                                                 |
+----------------------------------------------------------------------------------------------------------------------+
| The examination and testimony of the experts enabled the commission to conclude that five shots may have been fired. |
+----------------------------------------------------------------------------------------------------------------------+
```

### Type of change

- [x] Bug Fix (non-breaking change which fixes an issue)
- [x] New Feature (non-breaking change which adds functionality)
This commit is contained in:
Haruko386
2026-05-15 14:03:33 +08:00
committed by GitHub
parent 335dd5a263
commit c2863173b0
6 changed files with 483 additions and 16 deletions

View File

@@ -8,7 +8,9 @@
"models": "models",
"embedding": "embeddings",
"rerank": "rerank",
"balance": "user/info"
"balance": "user/info",
"tts": "audio/speech",
"asr": "audio/transcriptions"
},
"models": [
{
@@ -45,6 +47,27 @@
"model_types": [
"embedding"
]
},
{
"name": "fnlp/MOSS-TTSD-v0.5",
"max_tokens": 8192,
"model_types": [
"tts"
]
},
{
"name": "FunAudioLLM/CosyVoice2-0.5B",
"max_tokens": 8192,
"model_types": [
"tts"
]
},
{
"name": "FunAudioLLM/SenseVoiceSmall",
"max_tokens": 8192,
"model_types": [
"asr"
]
}
]
}

View File

@@ -5,7 +5,8 @@
},
"url_suffix": {
"chat": "chat/completions",
"models": "models"
"models": "models",
"tts": "audio/speech"
},
"class": "step",
"models": [
@@ -88,6 +89,27 @@
"chat",
"vision"
]
},
{
"name": "step-tts-2 ",
"max_tokens": 8192,
"model_types": [
"tts"
]
},
{
"name": "stepaudio-2.5-tts",
"max_tokens": 8192,
"model_types": [
"tts"
]
},
{
"name": "step-tts-mini",
"max_tokens": 8192,
"model_types": [
"tts"
]
}
]
}