IDEA-Research/grounding-dino-base
Primitive: /extract · Extract ·
Swin
The Grounding DINO model was proposed in Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection by Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, Lei Zhang.
Overview
Hardware: — drives latency, throughput & cost
| Size | 233M params |
|---|---|
| Tasks | /extract |
| License | apache-2.0 |
| Latency | 786 ms |
| Throughput | 0.8 mpix/s |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Extraction
| Output kinds | Bounding Boxes |
|---|---|
| Inputs | text · image |
| Max sequence length | — |
Benchmarks
COCO
Object detection on COCO natural images