Video-Text-to-Text
Transformers
Safetensors
English
videochat_flash_qwen
feature-extraction
multimodal
custom_code
Eval Results (legacy)
Instructions to use OpenGVLab/VideoChat-Flash-Qwen2-7B_res224 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use OpenGVLab/VideoChat-Flash-Qwen2-7B_res224 with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("OpenGVLab/VideoChat-Flash-Qwen2-7B_res224", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Update README.md
Browse files
README.md
CHANGED
|
@@ -93,8 +93,7 @@ VideoChat-Flash-7B is constructed upon UMT-L (300M) and Qwen2-7B, employing only
|
|
| 93 |
|
| 94 |
## 🚀 How to use the model
|
| 95 |
|
| 96 |
-
|
| 97 |
-
We provide the simple conversation process for using our model. You need to install [flash attention2](https://github.com/Dao-AILab/flash-attention) to use our visual encoder.
|
| 98 |
```
|
| 99 |
pip install transformers==4.39.2
|
| 100 |
pip install timm
|
|
|
|
| 93 |
|
| 94 |
## 🚀 How to use the model
|
| 95 |
|
| 96 |
+
First, you need to install [flash attention2](https://github.com/Dao-AILab/flash-attention) and some other modules. We provide a simple installation example below:
|
|
|
|
| 97 |
```
|
| 98 |
pip install transformers==4.39.2
|
| 99 |
pip install timm
|