Overview
mxbai-embed-large-v1 from Mixedbread AI scores 64.68 average on the MTEB benchmark across 56 tasks, placing it above OpenAI text-embedding-ada-002 (60.99) and BAAI BGE Large EN v1.5 (64.23). It produces 1024-dimensional L2-normalized vectors from inputs up to 512 tokens.
The model supports Matryoshka Representation Learning, meaning vectors can be truncated to 512 or 256 dimensions without retraining -- useful for cutting vector database storage costs on large corpora while retaining most retrieval quality.
Deploy it as a SageMaker endpoint in your own AWS account. Your documents stay inside your VPC -- no external API calls, no third-party access, no rate limits. You control the endpoint, the scaling policy, and the CloudWatch logs.
Integrates with LangChain, LlamaIndex, and Haystack. Compatible with pgvector on Aurora/RDS, Amazon OpenSearch, Pinecone, Weaviate, Qdrant, Chroma, and Milvus.
Primary use cases: highest-precision document retrieval for RAG pipelines, semantic search over internal knowledge bases, product catalog and patent similarity, and any workload where retrieval accuracy is the primary constraint.
Highlights
- MTEB score 64.68 -- beats OpenAI Ada-002 (60.99) and BGE Large (64.23), running entirely inside your AWS VPC
- Matryoshka support: truncate to 512 or 256 dimensions to cut vector DB storage without losing model quality
- Flat $0.10/hr on ml.m5.xlarge -- no per-token charges, no rate limits, 1024-dimensional CLS output
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Description | Cost/host/hour |
|---|---|---|
ml.m5.xlarge Inference (Real-Time) Recommended | Model inference on the ml.m5.xlarge instance type, real-time mode | $0.10 |
ml.m5.xlarge Inference (Batch) Recommended | Model inference on the ml.m5.xlarge instance type, batch mode | $0.10 |
Vendor refund policy
No refunds.
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
Amazon SageMaker model
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Version release notes
Initial release
Additional details
Inputs
- Summary
Mixedbread mxbai-embed-large-v1 on SageMaker. Scores 64.68 on MTEB -- beating OpenAI Ada-002 (60.99) and BGE Large (64.23). 1024-dimensional CLS embeddings, no per-token charges.
- Input MIME type
- application/json
Support
Vendor support
Contact support@waltsoft.net for deployment assistance.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.