# multimodal

このトピックのトレンドリポジトリ(4件)

bytedance/UI-TARS-desktop

bytedance/UI-TARS-desktopOtherTypeScript
36.6k3回登場

The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra

agentagent-tarsbrowser-usecomputer-usecoworkgui-agentgui-operatormcpmcp-servermultimodaltarsui-tarsvisionvlm

bojieli/ai-agent-book

bojieli/ai-agent-bookOtherPython
13.1k

《深入理解 AI Agent:设计原理与工程实践》(李博杰 著)开源主仓库:全书正文、编译版 PDF 与按章配套代码

agentagent-memoryai-agentbookcoding-agentcontext-engineeringlarge-language-modelsllmmcpmulti-agentmultimodalragreinforcement-learning

テキスト・画像・音声・動画をまるごと高速推論!万能AIモデルの配信基盤 — vllm-omni

vllm-project/vllm-omniAIPython
3.6k

vLLM-Omniは、テキストだけでなく画像・動画・音声など複数の種類のデータを同時に扱えるAIモデルを、高速かつ低コストで動かすためのフレームワーク(ソフトウェアの骨組み)です。もともとテキスト専用だったvLLMという人気の高速推論エンジ

audio-generationdiffusionimage-generationinferencemodel-servingmultimodalpytorchtransformervideo-generation

OpenMOSS/MOSS-TTS

OpenMOSS/MOSS-TTSOtherPython
2.7k2回登場

MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.

audioaudio-tokenizerllmmultimodaltext-to-speechvoice-cloning