Alibaba-NLP/gte-Qwen2-7B-instruct
Primitive: /encode · Encode ·
Qwen2
gte-Qwen2-7B-instruct is the latest model in the gte (General Text Embedding) model family that ranks No.1 in both English and Chinese evaluations on the Massive Text Embedding Benchmark MTEB benchmark (as of June 16, 2024).
Overview
Hardware: — drives latency, throughput & cost
| Size | 7.6B params |
|---|---|
| Tasks | /encode |
| License | apache-2.0 |
| Latency | 846 ms |
| Throughput | 3.5K tok/s |
| Cost | $0.063 /1M tok |
Cost is approximate — computed from list GPU prices; your actual price depends on the provider you deploy SIE with.
Embedding
| Output types | Dense |
|---|---|
| Dimensions | dense: 3,584 |
| Max sequence length | 32,000 |
| Inputs | text |
Benchmarks
NFCorpus
Biomedical literature search from NutritionFacts.org
NanoFiQA2018Retrieval
Smaller subset of the FiQA financial QA dataset