Skip to content
 
 

Latest commit

 

History

42 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Real-Gemini

Real-time video understanding and interaction through text,audio,image and video with large multi-modal model.

利用多模态大模型的实时视频理解和交互框架,通过文本、语音、图像和视频和这是世界进行问答和交流。

启动前端对话服务

主要实现了下面2个功能

  • 1、streamlit对话界面
  • 2、gpt4v请求接口

打开run.sh输入自己的api key, 然后启动

sh run.sh

TTS和ASR服务

  • ASR 服务调用自S组(TODO:要不要更新一个服务在这里)
  • TTS 见tts.py,启动脚本:
python tts.py

启动这些服务需要一些额外的环境和模型:torch, torchaudio, TTS,用pip安装即可,模型文件路径见py脚本。

Acknowledgement

关于我们 About Us

IDEA研究院封神榜团队是中文大模型开源计划Fengshenbang-LM的负责团队,开源包括二郎神系列太乙系列姜子牙系列等知名模型,并收获了开源社区的广泛使用和支持。

IDEA研究院CCNL技术团队已创建封神榜开源讨论群,我们将在讨论群中不定期更新发布封神榜新模型与系列文章。请扫描微信搜索“fengshenbang-lm”,添加封神空间小助手进群交流!

The IDEA Research Institute Fengshenbang team is the responsible team for the Chinese large model open source project Fengshenbang-LM. The open source includes well-known models such as the Erlang, Taiyi, and Ziya, and has received widespread use and support from the open source community.

The IDEA Research Institute CCNL technical team has created an open discussion group for Fengshenbang. We will periodically update and release new Fengshenbang models and series of articles in the discussion group. Please scan WeChat and search for "fengshenbang-lm", and add the Fengshen Space Assistant to join the group discussion!

About

Real-time video understanding and interaction through text,audio,image and video with large multi-modal model. 利用多模态大模型的实时视频理解和交互框架,通过文本、语音、图像和视频和这是世界进行问答和交流。

Topics

Resources

Stars

28 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages