Product Information
What is Ferret?
Apple's new multimodal large language model (MLLM) excels in both image comprehension and language processing, particularly demonstrating significant advantages in understanding spatial references.
How to use Ferret?
Ferret is Apple’s new multimodal large language model, excelling in image understanding and language processing, particularly in spatial referencing. It accurately identifies and describes spatial references in images with open vocabulary.
Core Functions of Ferret
Understand Spatial References of Any Shape and Granularity in Images
Accurate Localization with Open Vocabulary Descriptions
Accept Multiple Regional Inputs Like Points, Bounding Boxes, and Free Shapes
Excellent Performance in Region-based and Localization Multimodal Chats
Significantly Improved Image Detail Description Capability
Effectively Mitigates Object Hallucination Issues
Usage Scenarios of Ferret
- Performing Classic Reference and Positioning Tasks.
- Conducting Region-Based and Positioning-Required Multimodal Chats.
- Describing Image Content in Detail.
Common Questions about Ferret
What does Ferret do?
How do I use Ferret?
What are the core features of Ferret?
What are the use cases for Ferret?



















