Overview
Rerankers are neural networks that predict the relevancy scores between a query and documents and rank them based on the scores. They are used to refine search results in semantic search/retrieval systems and retrieval-augmented generation (RAG).
rerank-3-lite is the next generation of Voyage AI's reranker optimized for both latency and quality, and a drop-in upgrade to rerank-2.5-lite that requires no code changes. Trained with an updated backbone and an improved mixture of training data, rerank-3-lite improves on rerank-2.5-lite across nearly all domains, with the largest gains on long documents and code, and is roughly on par with the retrieval quality of rerank-2.5. Relevance scores are calibrated to match the score distribution of rerank-2.5-lite, so score thresholds tuned on rerank-2.5-lite continue to work.
Averaged across 95 retrieval datasets in 9 domains and four first-stage retrieval methods, rerank-3-lite outperforms Cohere Rerank v4.0 Pro, Qwen3-Reranker-8B, and rerank-2.5-lite by 2.23%, 2.53%, and 1.10% NDCG@10, respectively - outperforming Qwen3-Reranker-8B, the leading open-weights reranker, despite being over an order of magnitude smaller. On long-document retrieval, rerank-3-lite outperforms rerank-2.5-lite by 1.86% and Cohere Rerank v4.0 Pro by 13.80%. On code retrieval, it outperforms rerank-2.5-lite by more than 3.5% atop both voyage-3-large and voyage-4-large.
The model supports a combined context length of 32K tokens per query-document pair, including up to 8K tokens for the query, enabling more accurate retrieval over longer documents. rerank-3-lite also retains the instruction-following capability introduced in the rerank-2.5 series, allowing users to guide relevance scoring through natural language instructions. On the MAIR instruction-following benchmark, it outperforms rerank-2.5-lite by 0.91% and Cohere Rerank v4.0 Pro by 2.34%.
Learn more about rerank-3-lite here: https://blog.voyageai.com/2026/09/30/rerank-3
Highlights
- Optimized for both latency and quality, and a drop-in upgrade to rerank-2.5-lite with no code changes: same API, 32K-token combined context length per query-document pair (up to 8K for the query), and instruction following, with relevance scores calibrated so existing rerank-2.5-lite score thresholds continue to work.
- Outperforms Cohere Rerank v4.0 Pro by 2.23%, Qwen3-Reranker-8B by 2.53%, and rerank-2.5-lite by 1.10% NDCG@10 on average across 95 retrieval datasets spanning 9 domains, and is roughly on par with the retrieval quality of rerank-2.5.
- Largest gains on long documents and code: outperforms rerank-2.5-lite by 1.86% and Cohere Rerank v4.0 Pro by 13.80% on long-document retrieval, and rerank-2.5-lite by more than 3.5% on code retrieval atop both voyage-3-large and voyage-4-large.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Free trial
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.g6.xlarge Inference (Real-Time) Recommended | Model inference on the ml.g6.xlarge instance type, real-time mode | $2.25 |
ml.g5.xlarge Inference (Batch) Recommended | Model inference on the ml.g5.xlarge instance type, batch mode | $35.92 |
ml.g5.xlarge Inference (Real-Time) | Model inference on the ml.g5.xlarge instance type, real-time mode | $2.82 |
ml.g5.2xlarge Inference (Real-Time) | Model inference on the ml.g5.2xlarge instance type, real-time mode | $3.03 |
ml.g5.4xlarge Inference (Real-Time) | Model inference on the ml.g5.4xlarge instance type, real-time mode | $4.06 |
ml.g5.8xlarge Inference (Real-Time) | Model inference on the ml.g5.8xlarge instance type, real-time mode | $6.12 |
ml.g6.2xlarge Inference (Real-Time) | Model inference on the ml.g6.2xlarge instance type, real-time mode | $2.44 |
ml.g6.4xlarge Inference (Real-Time) | Model inference on the ml.g6.4xlarge instance type, real-time mode | $3.31 |
ml.g6.8xlarge Inference (Real-Time) | Model inference on the ml.g6.8xlarge instance type, real-time mode | $5.04 |
ml.g7e.2xlarge Inference (Real-Time) | Model inference on the ml.g7e.2xlarge instance type, real-time mode | $4.49 |
Vendor refund policy
Refunds are processed according to the EULA. For assistance, contact aws-marketplace@mongodb.com
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
MongoDB is excited to announce the initial release of rerank-3-lite.
Additional details
Inputs
- Summary
Supply a query and a list of documents to score for relevance.
Note: Does NOT support batch transform.
- Limitations for input type
- Total request size, calculated as the query tokens multiplied by the number of documents plus the tokens of all documents, cannot exceed 600,000 tokens.
- Input MIME type
- application/json
Input data descriptions
The following table describes supported input data fields for real-time inference and batch transform.
Field name | Description | Constraints | Required |
|---|---|---|---|
query | A query string (string). | Max 8,000 tokens. | Yes |
documents | Documents to rerank (List[string]). | Maximum of 1,000 documents; max 32,000 total tokens for each document. | Yes |
top_k | Number of most relevant documents to return (int). | Default: null (all documents are returned). | No |
truncation | Whether to truncate inputs to fit context limits (boolean). | Default: true. | No |
return_documents | Whether to return the documents in the response (boolean). | Default: false. | No |
Resources
Vendor resources
Support
Vendor support
Please email us at aws-marketplace@mongodb.com for inquiries and customer support.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Similar products

