google/owlv2-base-patch16-ensemble
Primitive: /extract · Extract ·
CLIP
The OWLv2 model (short for Open-World Localization) was proposed in Scaling Open-Vocabulary Object Detection by Matthias Minderer, Alexey Gritsenko, Neil Houlsby.
Overview
Hardware: — drives latency, throughput & cost
| Size | 155M params |
|---|---|
| Tasks | /extract |
| License | apache-2.0 |
| Latency | 955 ms |
| Throughput | 1.0 mpix/s |
| Cost | — /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Extraction
| Output kinds | Bounding Boxes |
|---|---|
| Inputs | image |
| Max sequence length | — |
Benchmarks
COCO
Object detection on COCO natural images