jina-vlm
A multilingual image-and-text model for visual questions and document analysis.
Compare models 0 / 3
Select two or three models to compare their published specifications.
- Released
- 2025-12-04
- Parameters
- 2.4B
- Context length
- 32K tokens
- Output dimensions
- Not published
- Inputs
- image · text
- Languages
- Multilingual
- Outputs
- text
- Availability
- api · license · huggingface · airgapped
Choose a representation
This model produces text or relevance rankings. A scalar embedding dimension does not describe its primary output.
Model weights and hosted API access
The license shown above applies to the published model assets. Calling a hosted API and downloading weights for commercial self-hosting are separate arrangements.
These weights are not listed under a permissive commercial license. Review the official license and obtain the required authorization before commercial self-hosting.
Read the official licenseCatalog checked on September 7, 2026. This is an independently prepared Chatsax guide to Jina models.