Listing Thumbnail

    Baseten

     Info
    Sold by: Baseten 
    Deployed on AWS
    Machine learning infrastructure that just works
    4.2

    Overview

    At Baseten, we provide all the infrastructure you need to deploy and serve ML models performantly, scalably, and cost-efficiently.

    With Baseten, you can:

    • Deploy your proprietary ML models with optimized serving engines.
    • Deploy open-source models on dedicated instances.
    • Handle massive traffic spikes with autoscaling model deployments.
    • Save on infra costs with scale to zero and lighting fast cold starts.
    • Manage deployments, metrics, and spending with role-based access control.

    Connect with us to discuss your ML infrastructure needs and learn more about our available live engineering support, custom POCs, volume discounts, and self-hosted options.

    Highlights

    • Highly performant autoscaling infrastructure that goes from prototype to production seamlessly.
    • Reliable logging and visibility across deployments, health, metrics, and spend in your Baseten workspace.
    • Enterprise-grade security and reliability with SOC 2 Type II, HIPAA compliance, and custom SLAs.

    Details

    Sold by

    Delivery method

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Trust Center

    Trust Center
    Access real-time vendor security and compliance information through their Trust Center powered by Drata or Vanta. Review certifications and security standards before purchase.

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    Pricing is based on the duration and terms of your contract with the vendor, and additional usage. You pay upfront or in installments according to your contract terms with the vendor. This entitles you to a specified quantity of use for the contract duration. Usage-based pricing is in effect for overages or additional usage not covered in the contract. These charges are applied on top of the contract price. If you choose not to renew or replace your contract before the contract end date, access to your entitlements will expire.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.

    1-month contract (1)

     Info
    Dimension
    Description
    Cost/month
    Baseten Base Package
    Listed pricing is indicative only. All purchases are completed via AWS Marketplace private offers tailored to your requirements. Reach out to request a custom quote and private offer
    $100,000.00

    Additional usage costs (1)

     Info

    The following dimensions are not included in the contract terms, which will be charged based on your usage.

    Dimension
    Description
    Cost/unit
    additional_usage
    Additional usage
    $1.00

    AI Insights

     Info

    Dimensions summary

    This listing uses a contract structure built around two dimensions. The Baseten Base Package covers your core committed spend, set through a private offer tailored to your needs. Additional usage bills separately for consumption beyond that base amount. You pay only for the compute your models actively use, billed by the minute, with no charge for idle time. All prices shown are indicative, and final terms come through a custom quote and private offer arranged with the vendor.

    Top-of-mind questions for buyers

    You pay only for the time your model actively uses compute, billed down to the minute. This covers deploying, scaling up or down, and making predictions. Idle time carries no charge. You control how your model scales up and down.
    The Baseten Base Package covers your committed spend arranged through a private offer. Additional usage bills separately for consumption beyond that base amount. Both appear together, with the base as your floor and additional usage capturing overage. Your active compute time drives what accrues.
    You can deploy open-source and custom models, or start from an off-the-shelf model library. Compute options include a range of GPU and CPU instance types, with control over which GPUs your models use. Reach out to the vendor to request additional GPU types or regions.
    www.baseten.co
    Helpful?

    Vendor refund policy

    All fees are non-refundable and non-cancellable except as required by law.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    Software as a Service (SaaS)

    SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.

    Support

    Vendor support

    Our standard email support is available Monday through Friday during business hours (Pacific time).

    We offer substantial additional support options, including Slack connect, live engineering support, custom POCs, and custom response SLAs.
    support@baseten.co 

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Product comparison

     Info
    Updated weekly
    By Baseten
    By Modal
    By Hugging Face

    Accolades

     Info
    Top
    10
    In Serverless Workloads
    Top
    10
    In High Performance Computing

    Customer reviews

     Info
    Sentiment is AI generated from actual customer reviews on AWS and G2
    Reviews
    Functionality
    Ease of use
    Customer service
    Cost effectiveness
    11 reviews
    Insufficient data
    14 reviews
    Insufficient data
    5 reviews
    Insufficient data
    Positive reviews
    Mixed reviews
    Negative reviews

    Overview

     Info
    AI generated from product descriptions
    Model Deployment and Serving
    Supports deployment of proprietary ML models with optimized serving engines and open-source models on dedicated instances.
    Autoscaling Infrastructure
    Handles massive traffic spikes with autoscaling model deployments and supports scale to zero functionality with fast cold starts.
    Monitoring and Observability
    Provides reliable logging and visibility across deployments, health metrics, and spending through a centralized workspace.
    Access Control and Management
    Includes role-based access control for managing deployments, metrics, and spending across the infrastructure.
    Security and Compliance
    Implements SOC 2 Type II certification, HIPAA compliance, and custom SLAs for enterprise-grade security and reliability.
    GPU Container Provisioning
    Custom infrastructure enables GPU-enabled container spin-up in approximately one second for rapid iteration and scaling.
    Autoscaling Capability
    Automatic scaling to hundreds of GPUs and down to zero resources within seconds without manual infrastructure configuration.
    Infrastructure-as-Code Deployment
    Python functions deployable to cloud using infrastructure-as-code approach with custom container image and hardware requirement definitions.
    Resource Optimization
    Dynamic resource allocation that scales up and down based on workload demands to optimize resource utilization.
    Serverless Compute Architecture
    Serverless platform supporting ML inference, fine-tuning, and batch data job execution without infrastructure management overhead.
    Model Deployment Infrastructure
    Inference Endpoints enable deployment of machine learning models as secure, production-ready APIs with fast inference capabilities.
    Application Hosting Platform
    Spaces provides hosting infrastructure for machine learning applications with integrated GPU resources and pre-configured dependencies.
    Enterprise Access Control
    Enterprise Hub includes Single Sign-On, Resource Groups, and Audit Logs for advanced security and access management.
    Model and Dataset Repository
    Platform hosts over 1 million pre-trained models, datasets, and AI applications for text, image, audio, and video processing tasks.

    Contract

     Info
    Standard contract

    Customer reviews

    Ratings and reviews

     Info
    4.2
    13 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    46%
    54%
    0%
    0%
    0%
    0 AWS reviews
    |
    13 external reviews
    External reviews are from G2 .
    KD pathak K.

    Reliable inference infrastructure for scaling custom open-source LLMs

    Reviewed on Oct 04, 2026
    Review provided by G2
    What do you like best about the product?
    It’s easy to deploy LLMs like Llama, Mistral, etc. on dedicated GPUs, and the API endpoint is clean and straightforward. I also like that I don’t need to manually manage a complex Kubernetes cluster to get everything running.
    What do you dislike about the product?
    Pricing doesn’t seem to scale well for heavy, enterprise-level GPU workloads, and I also wish the documentation were more granular—especially around custom Triton server setups.
    What problems is the product solving and how is that benefiting you?
    It helps reduce the infrastructure overhead of self-hosting machine learning models while also lowering inference latency for production AI applications.
    Computer & Network Security

    Remarkably Straightforward Model Packing and Shipping with Truss

    Reviewed on Oct 02, 2026
    Review provided by G2
    What do you like best about the product?
    Packing and shipping models with Truss is remarkably straightforward, especially compared with managing raw Kubernetes clusters or maintaining custom Triton setups.
    What do you dislike about the product?
    Debugging cryptic model packaging errors or dependency conflicts during container builds can take some trial and error in the CLI.
    What problems is the product solving and how is that benefiting you?
    Our team needed dedicated, high-throughput GPU infrastructure to run production ML models, without having to build an internal MLOps team from scratch.
    Najeeb K.

    Lightning-Fast Deployment with Cutting-Edge Truss Flexibility

    Reviewed on Oct 01, 2026
    Review provided by G2
    What do you like best about the product?
    I really appreciate how simple Truss makes deployment; we were able to move from a local model to a production endpoint in just a few hours thanks to the accurate documentation and examples. I also like how Baseten's autoscaling and optimized inference provided us with reliable, low-latency performance without overpaying for idle GPUs. Switching to Baseten from self-hosted GPUs and SageMaker offered us faster deployment and lower latency. Plus, not having to write Dockerfiles, handling dependencies, and model caching easily with Truss makes it even more useful for me.
    What do you dislike about the product?
    There's not really much that I have issues with, but if I had to pick something, it would be that some of the more advanced Truss configurations could use better documentation. Support helped us out quickly when we needed it, but it would be nice to have more end-to-end examples for things like private model weights with secrets and performance tuning, like concurrency and batching settings. The basic examples are great; it's just those edge cases that needed a bit more guidance.
    What problems is the product solving and how is that benefiting you?
    I use Baseten to deploy fine-tuned models, reducing latency and costs with its autoscaling. Truss simplifies deployment without needing Dockerfiles, supporting local testing and dependency handling.
    Prateek M.

    Baseten Makes ML Model Deployment Fast, Affordable, and User-Friendly

    Reviewed on Sep 29, 2026
    Review provided by G2
    What do you like best about the product?
    Mainly, I’ve used baseten to run and deploy ML models for testing and production deployment.

    There are a lot of things I really like about it, including the good UX, the speed, support documentation and is very cheap comparing other platforms.
    What do you dislike about the product?
    It doesn’t have strong downstream integration with other apps we use in our workflows. Also, the AI features feel quite limited compared with other platforms.
    What problems is the product solving and how is that benefiting you?
    It really helps simplify our deployment process and reduces the infrastructure work on our side. It also makes it easier to use the deployed models through APIs.

    Overall, this helps us build them more easily and then use them reliably later on.
    RAJ KUMAR U.

    Very simple and fast platform to deploy AI models

    Reviewed on Sep 26, 2026
    Review provided by G2
    What do you like best about the product?
    I really like how easy it is to deploy open-source models like DeepSeek. The dashboard is very clean and simple to use. Setting up my first model API took less than ten minutes, which saved me a lot of time
    What do you dislike about the product?
    The pricing can be a bit confusing for new users who are just starting out. Also, I noticed that the documentation for advanced model configurations is slightly hard to follow if You Don't Have a strong coding background
    What problems is the product solving and how is that benefiting you?
    It saves a lot of time because I do not have to worry about managing heavy cloud infrastructure or servers manually. I can just focus on testing different AI models quickly, which makes my development workflow much smoother
    View all reviews