Embed is a powerful tool that enables machines to comprehend and process text and images by converting them into numerical vectors. These vectors, with their 1024 dimensions, provide a comprehensive representation, allowing advanced AI applications to grasp the intricacies of user interactions, search queries, and document content.
One of its standout features is its multilingual capability, supporting over 100 languages. This enables cross-lingual searches, a valuable asset for global businesses and users seeking information across different language barriers.
Embed translates text and images into numerical vectors that models can understand. The most advanced generative AI apps rely on high-performing embedding models to understand the nuances of user inputs, search results, and documents. This Embed model has 1024 dimensions. This version is also a multilingual model that supports 100+ languages and can be used to search within a language (e.g., search with a French query on French documents) and across languages (e.g., search with a Chinese query on Finnish documents). Embed 3 is able to encode images to embeddings, images and texts must be encoded in separate requests. Added a new input type called images. Added a new parameter called images which accepts a data url of a base64 encoded png, jpeg, webp and gif, at this time, we only accept a single image per request.
Highlights
Cohere Embed is the leading multimodal (text and images) representation model used for semantic search, retrieval-augmented generation (RAG), classification, and clustering. As of September 2024, these models achieve state-of-the-art performance on a variety of text-to-image retrieval benchmarks in domains such as e commerce, UI/UX design and templates, and Graphs and Business Documents
Our optimized containers enable low latency inference on a diverse set of hardware accelerators available on AWS providing different cost and performance points for Bedrock customers.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
This listing bills the multilingual embedding model two ways. On-demand usage charges you per 1 million input tokens, per 1 million tokens, or per image processed, so cost follows the volume you send. If you need steady capacity, provisioned throughput charges an hourly rate per instance instead. Provisioned throughput comes in three commitment levels: no commit, 1-month commit, and 6-month commit. Longer commitments set the hourly rate for reserved capacity. You pick usage-based pricing for variable workloads or provisioned throughput for predictable, dedicated capacity.
Top-of-mind questions for buyers
What counts as one input token for the per-token pricing?
A token is a chunk of text the model processes, not a character or word. Simple text averages about one token per word. Complex text with uncommon words may use three to four tokens per word. Token counts drive your per-million-token charges.
How does the per-token option differ from provisioned throughput for billing?
Per-token and per-image pricing meters the actual volume you send, so charges follow real usage with no commitment. Provisioned throughput charges a fixed hourly rate per instance for reserved capacity, whether or not you use it. Usage-based suits variable workloads; provisioned throughput suits steady, predictable volume.
Am I charged when a provisioned throughput instance sits idle?
Provisioned throughput bills a fixed hourly rate per instance for reserved capacity. The charge applies for the reserved time regardless of how many tokens or images you process. Idle capacity still accrues the hourly rate, so you pay for the dedicated instance, not the volume run through it.
Request a private offer to receive a custom quote.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Connect your OpenAI, Anthropic, and Cohere applications to Amazon Bedrock by changing one base URL. 100+ models, plus images, audio, files and embeddings. Runs in your own AWS account.
Deploy a private LLM inference server in your VPC in minutes. Pre-integrated Ollama, GPU stack, RAG pipeline, and 20+ AI frameworks - no external API calls.
Business Compass AI Chatbot is a white-label, RAG-powered assistant built on Amazon Bedrock that answers questions from your approved content only. Deploys into your AWS account.
Expert deployment of stdapi.ai, the OpenAI, Anthropic and Cohere AI gateway for Amazon Bedrock, into your own AWS account. Production infrastructure configured and secured by specialists.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.