Listing Thumbnail

    Cohere Rerank v3.5 (Amazon Bedrock Edition)

     Info
    Sold by: Cohere 
    Deployed on AWS
    AWS Free Tier
    Rerank improves search systems by sorting documents based on their semantic similarity to a query.
    4.5

    Overview

    Rerank v3.5 endpoint enables businesses to significantly improve search and retrieval-augmented generation systems. As input, it takes a query and list of potentially relevant documents. Rerank v3.5 then returns the documents as a list sorted by semantic similarity to the provided query. As an intelligent cross-encoding AI model, Rerank v3.5 is able to understand the meaning behind enterprise data and user questions. Rerank v3.5 can be implemented with just a few lines of code, delivers leading performance across over 100 languages, and is uniquely capable of understanding complex information which requires reasoning. These attributes make Rerank v3.5 particularly well suited for global organizations within Finance, Healthcare, Energy, Government, and Manufacturing. Rerank v3.5 can be added to existing systems, whether keyword or semantic, to improve performance.

    Highlights

    • Cohere's Rerank v3.5 is uniquely capable of understanding complex documents and queries. This leads to more accurate search results when user questions have multiple aspects and require reasoning. Rerank v3.5 also offers strong performance on semi-structured data such as Code, Tables, and JSON Documents. These attributes make the model ideal for global organizations within such as Finance, Healthcare, Energy, Government, Manufacturing.
    • Cohere's Rerank v3.5 can be added to existing search and retrieval-augmented generation (RAG) systems with just a few lines of code. This ease of implementation makes is simple to boost semantic understanding and improve search results. Rerank v3.5 is also highly efficient, in terms of throughput, and is capable of satisfying demanding requirements for large organizations.
    • Cohere's Rerank v3.5 offers leading multilingual performance in over 100 languages, including but not limited to: Arabic, Chinese, English, French, German, Hindi, Japanese, Korean, Portuguese, Russian, and Spanish. This is useful for global organizations who operate across various languages and require a performant AI model to improve their search systems.

    Details

    Sold by

    Delivery method

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    Cohere Rerank v3.5 (Amazon Bedrock Edition)

     Info
    Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.

    Usage costs (5)

     Info
    Dimension
    Cost/unit
    Price per 1 search unit
    $0.002
    Price per 1 search unit
    $0.002
    Price per 1 search unit
    $0.002
    Price per 1 search unit
    $0.002
    Price per 1 search unit
    $0.002

    AI Insights

     Info

    Dimensions summary

    You pay for this reranking model based on usage, measured in search units. One search unit equals one query with up to 100 documents to rank. If a document exceeds 500 tokens, including your query, it splits into chunks. Each chunk counts as a separate document toward the total ranked. All five dimensions bill by the same search unit and share this pricing logic. Your cost scales with the number of searches you run and the number of documents each search ranks.

    Top-of-mind questions for buyers

    If a document exceeds 500 tokens, including your query length, it splits into multiple chunks. Each chunk counts as an individual document toward the 100-document limit for that search unit. Longer documents therefore consume more of your document allowance per search.
    One search unit covers one query with up to 100 documents. Ranking beyond 100 documents in a query goes past a single search unit's limit. Your usage scales with both the number of queries and how many documents each query ranks.
    You pay on a usage basis for each search unit processed. Charges accrue as you run queries and rank documents. There is no charge when you send no requests, since billing meters actual search units rather than a fixed period.
    cohere.com
    Helpful?

    Vendor refund policy

    No refunds. Please contact support+aws@cohere.com  for further assistance.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    Software as a Service (SaaS)

    SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.

    Support

    Vendor support

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Similar products

    Customer reviews

    Ratings and reviews

     Info
    4.5
    1 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    100%
    0%
    0%
    0%
    0%
    1 AWS reviews
    Singh Aman

    RAG assistant workflows have improved answer relevance and now deliver faster accurate decisions

    Reviewed on May 23, 2026
    Review from a verified AWS customer

    What is our primary use case?

    My main use case for Cohere Rerank v3.5 is that we used it in our enterprise RAG-based AI assistant solution hosted over AWS, which was integrated with Amazon Bedrock, OpenSearch, and vector embeddings. To improve document retrieval relevance for internal enterprise workflows, such as access management, ticketing, and knowledge retrieval, we have used Cohere Rerank v3.5.

    A specific example of how we used Cohere Rerank v3.5 in one of those workflows is that we usually found some bugs when asking questions as the KB search was doing multiple calls. To avoid that, we thought we would use a reranking model to decrease the KB calls, ensuring we have proper KBs at the first go, and for that, we have used reranking.

    What is most valuable?

    The best features Cohere Rerank v3.5 offers include high-quality reranking relevance, better contextual retrieval, easy AWS integration, fast response time, and strong multilingual support, which are some of the better things that I have observed.

    The easy AWS integration and fast response time help my team because the whole solution is on AWS, and AWS already provides Cohere as a provider where reranking is available. We just have to call the ARN of the reranking model as an interface, and it easily integrates into the KB search call, making integration straightforward.

    The most valuable feature was the reranking quality. After introducing Cohere Rerank v3.5 into our pipeline, the relevance of the required chunks improved significantly, which directly reduced hallucination responses from the downstream LLMs, and the latency was quite good, making it acceptable for the enterprise-grade application.

    Cohere Rerank v3.5 has positively impacted my organization by improving answer accuracy in our AI assistant workflows and reducing irrelevant retrieval results. This improved end-user trust in the system and helped move some proof of concept implementation closer to production readiness so that our end users can trust the answers.

    We have significantly seen the outcomes, and the answer quality outcomes have improved after implementing the reranking.

    What needs improvement?

    For improvement purposes, latency can be improved for sure, as it is currently around one to one and a half seconds, and if we can improve it so that it takes much lesser time than whatever it is taking right now, that would be great. Other than that, I have not seen that much room for improvement because it is already a much improved version I am using right now.

    I can see that better native observability can be implemented, and price transparency is not there on AWS. Other than that, AWS native analytics could also be helpful for developers, and if fine-tuning can be available for those reranking models, it could have much better control over the reranking model, in my opinion.

    I do not think there are any other improvements for Cohere Rerank v3.5 that we have not discussed yet, as we have already talked about latency, integration, and everything else that is already there.

    For how long have I used the solution?

    I have been using Cohere Rerank v3.5 for the last one and a half years.

    What do I think about the stability of the solution?

    Cohere Rerank v3.5 is quite stable and fast.

    What do I think about the scalability of the solution?

    Cohere Rerank v3.5's scalability is something that works on-the-go, as AWS already supports scalability for enterprise-specific needs. If right now one hundred users are using it, fewer resources will be utilized, but if more than that or maybe one thousand to ten thousand users are using it, the load will scale accordingly, and we have not seen any performance degradation as user numbers increase, so scaling works very fast and much better.

    How are customer service and support?

    I have not yet visited the customer support for Cohere Rerank v3.5 because we have not required that. In terms of stability and scalability, we have found that the solution was stable during testing in enterprise workloads and was able to handle large document retrieval scenarios with acceptable performance degradation. The API integration through AWS was straightforward and reliable as all of this was mentioned in the AWS documentation, which was quite good.

    Which solution did I use previously and why did I switch?

    We were not initially using any reranking models, but after switching to reranking models, the performance and answer quality have improved significantly.

    How was the initial setup?

    The setup process is very easy as the inference model is already provided on the AWS documentation on how to utilize it, which I think is very good to have.

    What was our ROI?

    I have seen a return on investment, as time has been saved significantly because the end-user experience has improved considerably, and the end-user is impressed with the responses we are providing to them, appreciating the response quality greatly.

    Which other solutions did I evaluate?

    Before choosing Cohere Rerank v3.5, I evaluated other options, including AWS Titan, which provides embedding retrieval, as well as OpenSearch k-NN and some open-source reranking models from Hugging Face, but we found that Cohere performs much better than the options I mentioned earlier.

    What other advice do I have?

    I would suggest to any developer who wants to increase their RAG response quality to look into Cohere Rerank v3.5 for a significant improvement in response quality. I gave this product a rating of nine out of ten.

    View all reviews