Collections
Discover the best community collections!
Collections including paper arxiv:2303.15343
-
PaliGemma: A versatile 3B VLM for transfer
Paper • 2407.07726 • Published • 73 -
Vision language models are blind
Paper • 2407.06581 • Published • 84 -
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Paper • 2404.16994 • Published • 39 -
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Paper • 2403.05525 • Published • 49
-
Clinical Text Summarization: Adapting Large Language Models Can Outperform Human Experts
Paper • 2309.07430 • Published • 27 -
MindAgent: Emergent Gaming Interaction
Paper • 2309.09971 • Published • 12 -
Cure the headache of Transformers via Collinear Constrained Attention
Paper • 2309.08646 • Published • 13 -
Contrastive Decoding Improves Reasoning in Large Language Models
Paper • 2309.09117 • Published • 39
-
NVLM: Open Frontier-Class Multimodal LLMs
Paper • 2409.11402 • Published • 75 -
BRAVE: Broadening the visual encoding of vision-language models
Paper • 2404.07204 • Published • 20 -
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Paper • 2403.18814 • Published • 49 -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Paper • 2409.17146 • Published • 123
-
Sigmoid Loss for Language Image Pre-Training
Paper • 2303.15343 • Published • 12 -
google/siglip-base-patch16-224
Zero-Shot Image Classification • 0.2B • Updated • 1.85M • 92 -
google/siglip-base-patch16-256
Zero-Shot Image Classification • 0.2B • Updated • 75.1k • 8 -
google/siglip-base-patch16-384
Zero-Shot Image Classification • 0.2B • Updated • 68.6k • 12
-
NVLM: Open Frontier-Class Multimodal LLMs
Paper • 2409.11402 • Published • 75 -
BRAVE: Broadening the visual encoding of vision-language models
Paper • 2404.07204 • Published • 20 -
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Paper • 2403.18814 • Published • 49 -
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models
Paper • 2409.17146 • Published • 123
-
PaliGemma: A versatile 3B VLM for transfer
Paper • 2407.07726 • Published • 73 -
Vision language models are blind
Paper • 2407.06581 • Published • 84 -
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Paper • 2404.16994 • Published • 39 -
DeepSeek-VL: Towards Real-World Vision-Language Understanding
Paper • 2403.05525 • Published • 49
-
Sigmoid Loss for Language Image Pre-Training
Paper • 2303.15343 • Published • 12 -
google/siglip-base-patch16-224
Zero-Shot Image Classification • 0.2B • Updated • 1.85M • 92 -
google/siglip-base-patch16-256
Zero-Shot Image Classification • 0.2B • Updated • 75.1k • 8 -
google/siglip-base-patch16-384
Zero-Shot Image Classification • 0.2B • Updated • 68.6k • 12
-
Clinical Text Summarization: Adapting Large Language Models Can Outperform Human Experts
Paper • 2309.07430 • Published • 27 -
MindAgent: Emergent Gaming Interaction
Paper • 2309.09971 • Published • 12 -
Cure the headache of Transformers via Collinear Constrained Attention
Paper • 2309.08646 • Published • 13 -
Contrastive Decoding Improves Reasoning in Large Language Models
Paper • 2309.09117 • Published • 39