Overview
MS-MARCO MiniLM-L-12-v2 is the 12-layer variant of the MS-MARCO cross-encoder family with 2.9 million monthly downloads, delivering meaningfully higher reranking accuracy than the L-6 variant at moderate throughput. With 12 transformer layers processing query-document pairs jointly, it captures deeper relevance signals for complex enterprise queries -- technical documentation, legal language, financial terminology, and multi-part questions where shallow models miss nuance. At 134MB it remains fully CPU-native and deployable without GPU infrastructure. Use it when retrieval quality directly affects user outcomes: enterprise knowledge management, compliance document search, and contract review assistance where extra precision justifies slightly higher latency.
Highlights
- 12-layer architecture delivers higher reranking precision than L-6 for complex enterprise queries
- 2.9M monthly downloads; the quality choice when RAG accuracy determines business outcomes
- 134MB fully CPU-native -- no GPU needed, deploys on ml.m5.xlarge
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.m5.xlarge Inference (Real-Time) Recommended | Model inference on the ml.m5.xlarge instance type, real-time mode | $0.10 |
ml.m5.xlarge Inference (Batch) Recommended | Model inference on the ml.m5.xlarge instance type, batch mode | $0.10 |
Vendor refund policy
No refunds.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Initial release
Additional details
Inputs
- Summary
2.9M monthly downloads. Deeper 12-layer cross-encoder reranker for higher-accuracy RAG. Optimal precision-throughput balance for enterprise search.
- Input MIME type
- application/json
Support
Vendor support
Contact support@waltsoft.net for deployment assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.