Automated testing for Data Engineers. Datafold is the fastest way to validate dbt model changes during development, deployment & migrations. With Datafold, data engineers can audit their work in minutes without writing tests or custom queries. Integrated into CI, Datafold enables data teams to deploy with full confidence, ship faster, and leave tedious QA and firefighting behind.
Datafold is the fastest way to validate dbt model changes during development, deployment & migrations. With Datafold, data engineers can audit their work in minutes without writing tests or custom queries. Integrated into CI, Datafold enables data teams to deploy with full confidence, ship faster, and leave tedious QA and firefighting behind.
Automate proactive testing for all data transformations. Supercharge your dbt workflows with seamless dbt Cloud and Core integrations.
Know exactly what will happen to data and data applications once the code is deployed, right in the pull request. Identify breaking changes, sudden metric shifts and edge cases before they do any damage to the business.
Stop guessing what this regex does or arguing if that CASE WHEN statement has correct logic. No more custom scripts and audit spreadsheets to fill.
Stop surprising your data users with unexpected metric changes and broken dashboards. Easily share impact reports with everyone and give heads up before deploying the changes to production.
Manual data testing is hard, tedious, and error-prone. Focus on what matters and not on writing boilerplate tests, custom scripts and filling out audit spreadsheets.
With full visibility into every change, everyone, not just data team, can contribute, because testing and reviewing code is so easy!
Access real-time vendor security and compliance information through their Trust Center powered by Drata or Vanta. Review certifications and security standards before purchase.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor, and additional usage. You pay upfront or in installments according to your contract terms with the vendor. This entitles you to a specified quantity of use for the contract duration. Usage-based pricing is in effect for overages or additional usage not covered in the contract. These charges are applied on top of the contract price. If you choose not to renew or replace your contract before the contract end date, access to your entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
Datafold prices this listing by the number of provisioned developer seats under a contract. You choose between two options: one for 5 developers and one for 10 developers. Both bill as a fixed commitment for the users you provision, not by usage. The difference is seat count alone — you pick the tier that matches your team size. To scale, you move to the option with more seats. Pricing scales with the number of developers you need to cover.
Top-of-mind questions for buyers
What counts as one provisioned developer for billing purposes?
A provisioned developer is a user seat you set aside in advance. You pay for the seats you provision, not for how much each person uses the platform. The 5 and 10 developer options simply cover 5 or 10 such seats regardless of daily activity.
What happens if my team grows beyond the developer seats I purchased?
Seat count is a fixed commitment for the term. If your team outgrows the 5 developer option, you move to the 10 developer option to cover more users. The seats you provision set your cost; adding people means selecting a plan with a higher seat count.
Is this billed by usage or as a fixed commitment?
This is a contract commitment based on provisioned seats, not usage. You commit to a set number of developers and pay a fixed amount for the term. Your bill does not change with how much data you diff, monitor, or reconcile — only with the seat count you choose.
www.datafold.com
Helpful?
Vendor refund policy
All fees are non-cancellable and non-refundable except as required by law.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Seamless integration with dbt Cloud and Core for automated testing of data transformations and model changes
Automated Data Validation
Automated testing capability that validates data model changes without requiring manual test writing or custom queries
CI/CD Pipeline Integration
Integration into continuous integration workflows to enable automated validation during development, deployment, and migration phases
Change Impact Analysis
Detection and identification of breaking changes, metric shifts, and edge cases in data transformations before production deployment
Data Quality Reporting
Generation of impact reports and visibility into data changes for stakeholder communication and deployment decision-making
AI-Native Code Development Environment
Integrated development environment with AI capabilities for coding data pipelines using dbt and Python, featuring built-in warehouse access and column-level lineage context.
State-Aware Pipeline Scheduling
Scheduler supporting state-aware execution of dbt and Python data pipelines with column-level impact analysis for CI testing.
Data Lineage and Impact Analysis
Column-level lineage tracking and impact analysis capabilities for understanding data dependencies and transformation effects across pipelines.
Warehouse Cost Optimization
AI-agent based monitoring and optimization system operating continuously to reduce warehouse operational costs.
Data Pipeline Orchestration Integration
Support for orchestrating multi-tool data workflows including Fivetran ingestion, data transformation pipelines, and downstream application refreshes for Tableau and PowerBI.
Integrated Development Environment
IDE built for dbt with SQL Runner capable of executing Jinja templates and guided version control enforcement for git best practices
Job Scheduling and Orchestration
Custom scheduling capabilities for production jobs with incremental testing triggered on change or before deployment
CI/CD Pipeline Support
Continuous integration and continuous deployment functionality for automated dbt project workflows
Enterprise Security and Compliance
SOC2 Type II compliance certification, single sign-on (SSO) authentication, and role-based access control
Monitoring and Alerting
Built-in monitoring and alerting capabilities for tracking job execution and system health
Data monitoring has transformed our secure data workflows and now drives accurate decisions
Reviewed on Aug 10, 2026
Review provided by PeerSpot
What is our primary use case?
My main use case for Datafold is as a data observability platform that we leverage for many of our data products, particularly because our data is very susceptible to any kind of breakdown or catastrophe due to its confidential nature. Datafold helps us identify, prioritize, and deep dive into data quality issues which can occur anytime. The platform finds these issues and investigates them in a proactive manner, and apprises us to rectify them before they hit production. Currently, we are using it to improve the quality of our organizational assessments and tests. It helps us test our various products and services, validate various parameters, and monitor whether the production cycle is happening in a synchronous manner. It is assisting us significantly in terms of migrating data from less secure platforms to more secure platforms as per clientele need. Datafold is continuously working 24/7 contextualizing the data ensuring that migrations happen accurately, as well as reconciling data to ensure there are no gaps, providing complete monitoring from start to end.
Our organization started deploying Datafold primarily due to outstanding peer feedback we received regarding its user-friendly interface and the quality of data monitoring, which is amazing. Since we deal in agriculture, where personal farmer data must not leak, Datafold helps us simplify the complicated process of collecting diverse data in a streamlined manner and track it from start to end, ensuring a high level of accuracy is maintained during data collection, transferring, migrating, and ultimately monitoring. This has significantly improved our operations, and we are continuing with Datafold.
What is most valuable?
Regarding the best features Datafold offers, tracking and monitoring of our data pipeline has become easier than ever, especially in the field of agriculture where collecting vast amounts of data is critical. India, being an agrarian country, has a majority of its population dependent on agriculture, making the tracking of complex data essential. The user interface is very useful and easy to navigate, allowing newcomers to train for just a week before being able to work independently with Datafold. It simplifies the entire process of collecting data from source to sync and addresses concerns related to secret and confidential data by supporting integration effectively. Datafold has consistently proven itself with its real-time alerts and visualizations when analyzing data for insights, enabling users to check on real-time anomalies and resolve them quickly. From our perspective, the platform has improved our testing environment significantly, ensuring the quality and consistency of captured data while saving us time and effort from manual intervention.
I find Datafold to be superior, given its strong positive feedback from numerous users on platforms such as Google. My team has validated that the data testing capabilities are impressive, helping users validate data quality and identify any issues or lag, thus enabling us to address root causes before they become significant problems. Datafold has proven to be a game changer for organizations, and its pricing is quite reasonable, making it easy on our budget and encouraging annual subscription renewals.
Datafold has positively impacted my organization as the user interface is extremely easy to navigate, allowing my team to efficiently engage with the environment and receive real-time updates regarding data capturing, quality, migration, and flow. It alerts us to any bugs, enabling us to reduce or eliminate errors almost entirely. Datafold provides an environment for creating tests, allowing users to independently ensure data matches benchmark quality and consistency, which saves valuable manpower and time. Thanks to Datafold, our team is now focused on more critical tasks rather than debugging and monitoring data flow. Overall, the feedback regarding data handling is consistently very positive.
In terms of measurable impact, when we used traditional methods, our team consumed approximately four to five man hours daily for data transfers, but with Datafold, we have reduced that to only one hour per day, which is a substantial time-saving and a game changer for our organization.
What needs improvement?
Currently, I do not have specific feature requests or concerns since I have not heard of any major issues. A few limited integration issues arose during our initial two years but were resolved promptly, and there are no lingering drawbacks as of now. While we sometimes feel that more features could be added to keep pace with the evolving data migration environment, I do not see any significant drawbacks. Ideally, I would like to see an expansion of features as the platform has great potential.
For how long have I used the solution?
I have been using Datafold in the current organization for the past two years.
What do I think about the stability of the solution?
Datafold is absolutely stable, which is why we have used it for two years. It has brought significant business stability by ensuring data quality, monitoring, migration, and flow are structured and streamlined. The platform is reliable, with no data breakage or lag in operations, making it highly scalable from when we began with 40 to 50 clients to now over 200 clients without issues. Datafold provides powerful capabilities for automated data validation with no manual intervention, ensuring robust data security and preservation.
What do I think about the scalability of the solution?
Datafold is absolutely stable, which is why we have used it for two years. It has brought significant business stability by ensuring data quality, monitoring, migration, and flow are structured and streamlined. The platform is reliable, with no data breakage or lag in operations, making it highly scalable from when we began with 40 to 50 clients to now over 200 clients without issues. Datafold provides powerful capabilities for automated data validation with no manual intervention, ensuring robust data security and preservation.
How are customer service and support?
Customer support is excellent, available 24/7 through various channels, including phones, emails, and a ticketing system. I would rate customer support as a four out of five as there are two types of customer support: one for general feedback and another technical team addressing integration issues. Overall, it is a great support system.
Which solution did I use previously and why did I switch?
We have not used any other solutions long-term as the homegrown options developed by our tech team were not effective, leading us to switch to Datafold immediately.
How was the initial setup?
In terms of pricing, setup costs, and licensing, I find everything to be reasonable, quite affordable, and stable over the years as the prices have not increased more than ten percent in two years. The initial licensing and setup were also straightforward, and while I cannot disclose specifics about the quote, it was definitely negotiable. Implementation and integration costs were minimal, allowing small and marginal companies to consider Datafold.
What was our ROI?
There has been a clear return on investment since the first year because we have been able to generate value from Datafold's implementation. Previously, eight people worked on data processes, but now only two are required, significantly reducing time, costs, and human resources, all while enhancing our image in front of clients. Overall, we have no issues with data errors or migration failures, showcasing all the positives that contribute to our ROI.
What's my experience with pricing, setup cost, and licensing?
In terms of pricing, setup costs, and licensing, I find everything to be reasonable, quite affordable, and stable over the years as the prices have not increased more than ten percent in two years. The initial licensing and setup were also straightforward, and while I cannot disclose specifics about the quote, it was definitely negotiable. Implementation and integration costs were minimal, allowing small and marginal companies to consider Datafold.
Which other solutions did I evaluate?
We did not evaluate many options before choosing Datafold as our experience during the trial version was so positive that we decided to stick with it instead of exploring others.
What other advice do I have?
For others considering Datafold, I advise them to assess their own needs before committing to the platform. Datafold can be a game changer for handling vast data, data migration, and monitoring. It is important for organizations to clearly define their ROI metrics relevant to their context and establish clear, understandable SLAs that their team can implement. My final thoughts about Datafold are that any size organization would benefit from its ability to manage complex data infrastructure and architecture effortlessly, providing a seamless data engineering and automation experience. I would rate this solution a nine out of ten.
Caio Gabriel Guimarães
Data quality checks have become streamlined and validate complex migration transformations
Reviewed on Jul 29, 2026
Review provided by PeerSpot
What is our primary use case?
I used Datafold for data quality checks while working on a data migration project where we moved data and applied business rules and transformations. After that movement, we had to check if the data loaded was correct, and Datafold helped significantly because it has great resources for customizing queries and defining primary keys for comparison. We could set the source and target tables, specify rules, and then press play, which provided excellent results regarding data matching percentages.
What is most valuable?
I used Datafold for data quality checks while working on a data migration project where we moved data and applied business rules and transformations. After that movement, we had to check if the data loaded was correct, and Datafold helped significantly because it has great resources for customizing queries and defining primary keys for comparison. We could set the source and target tables, specify rules, and then press play, which provided excellent results regarding data matching percentages.
Sometimes the mismatches were minor, but other times they required more attention. Being able to use Datafold for these quality checks while working on other processes was really helpful.
What needs improvement?
I know that bugs can be related to misconfigurations, but we had issues with comparisons where executing the queries simply did nothing, and we didn't have much information about why it failed. I believe having more detailed information about why the comparison didn't work would help with debugging, so a more detailed log would be really beneficial.
For how long have I used the solution?
I worked with Datafold for one project that took about one year, and that is my experience with it so far.
What do I think about the stability of the solution?
We had experiences where processes took longer than expected, but it was unclear if it was related to Datafold or the database systems. Overall, it performed well, but sometimes not specifying the number of rows for comparison led to execution issues, as the tool struggled with high data volumes, which required us to find the right amount of data for comparison.
What do I think about the scalability of the solution?
I wasn't involved in scalability issues, and I don't think I had access to those configurations as a data engineer, focusing instead on data quality checks without delving into setup or configuration.
How are customer service and support?
I haven't contacted technical support or customer support from Datafold during the time I worked with it.
Which solution did I use previously and why did I switch?
Before working with Datafold, I have never used similar tools. I know there are some options, but I haven't worked with them. My experience required writing customized queries or automation scripts for tasks that Datafold automates.
I didn't work with a similar tool before or after Datafold. Whenever I needed to do data quality checks, I always had to write a customized query or script.
How was the initial setup?
The initial deployment was really easy. Once we understood how it worked and set it up, connecting to databases like Redshift, SQL Server, and Databricks was straightforward.
What's my experience with pricing, setup cost, and licensing?
I'm not familiar with the pricing details, as it's not information that comes to us as data engineers, but I know it has a high cost. The client I worked with raised concerns about the pricing, but I'm not involved with the project anymore, so I don't know if they're still using it.
What other advice do I have?
I talked with someone on LinkedIn about PeerSpot, and he mentioned something about a gift card to do the review, which is why I asked for this meeting because I'm not comfortable using the company email and I'm not sure if I'm authorized to do that. He explained that I should do a review about some tool that I have experience with to be eligible for a possible gift card, but I didn't know about the company before, to be honest.
I have been working in my current field for 11 years overall, dealing with data-related projects.
I have no questions and will be waiting for the next steps in the process. My review rating for Datafold is 8.
JohnBosco Obi
Data diffing has caught regressions early and now reporting and usability still need improvement
Reviewed on Jul 18, 2026
Review provided by PeerSpot
What is our primary use case?
My main use case for Datafold is to automatically compare data differences through data diffing, where I can compare datasets row-by-row, column-by-column, to surface unexpected changes.
I also use it for CI/CD integration for data pipelines, where we embed data quality checks into GitHub and GitLab pull requests. Another interesting feature I use Datafold for is column-level lineage to trace how specific columns flow through SQL transformations across the data stack we have.
Recently, for data diffing, there was one time we wanted to conduct a lot of comparisons on a very large dataset to be sure that all was in check before we migrated to Snowflake, which is a complex and resource-heavy process.
We used Datafold to do the automatic checks across rows and columns to ensure that when we completed the migration, no data was lost and the migration into Snowflake was very smooth. That was in March, and it was very useful.
For a friend who uses Datafold in their enterprise, an engineer made a change from an SQL model and did not realize it would cascade and alter downstream dashboard metrics.
However, with the help of Datafold integrated into the pull request workflow, the data diff actually surfaced exactly which columns changed, which rows were affected, and by how much.
What is most valuable?
The data diff stands out the most for me as the best feature because it gives me that column-level and row-level comparison between any two datasets, highlighting what changed, even down to characters and whitespaces.
It also tells me when something has failed, and it makes debugging faster and more precise.
Column-level lineage derived from SQL static analysis helps me trace how columns flow through transformations across the entire pipeline.
Datafold has positively impacted my organization by helping us catch data regressions before they reach production.
Even while we are still building the pipeline and mapping out how the data will flow, we catch any data regressions and alterations before they reach production. This helps us avoid a lot of debugging when the migration has happened or when we reach production level and data goes live.
Additionally, it significantly reduces manual testing time, saving us time that we can put into useful work. Furthermore, it builds our team's confidence when pushing code changes, as the data reviewers, data annotators, and engineers can actually see the data impact of every change, not just the code change.
What needs improvement?
The first pain point for me is that the reporting capabilities are weak. The ease of setup is also challenging, particularly for those who are not tech-savvy or do not know how to navigate it. Most importantly, there is no free trial, so you cannot deploy and test it to see the efficiency and use case before purchasing.
I feel the features need more attention because sometimes finding the data I want to use or compare is challenging, especially when looking for it inside the SaaS.
Automated workflows also break sometimes, but whenever that happens, the good thing is that it gives you an error report so you know where the issue is coming from.
For how long have I used the solution?
I have been using Datafold for about six months.
What other advice do I have?
I give Datafold a seven out of ten because it is excellent at its core specialization, which is data diffing and CI/CD integration, and its migration automation capabilities are strong and very AI-powered. However, I lose three points due to the weak reporting.
Even though it provides reports when there is a breakage in your push or migration, I sometimes cannot get the full scope of what I want when producing an actual report. Additionally, the lack of a free trial is a downside.
Regarding Datafold's governance and security, I rate it high because its migration agent uses LLMs for SQL translation and validation, which is a strong point.
Additionally, there is a self-hosted deployment option available, so organizations with strict data residency or compliance requirements can run Datafold entirely within their own cloud environment, which could be either AWS, GCP, or Azure, ensuring that data never leaves the perimeter of the organization.
Datafold's output is highly accurate and very reliable. The migration agent's accuracy, which utilizes LLMs to convert SQL dialects, is excellent.
Furthermore, the data diff, which is the most reliable AI-adjacent feature, is deterministic, not generative, and it compares actual data values mathematically rather than using inference.
Because it does the comparison mathematically, the outputs are highly accurate and consistently reliable.
Datafold is deployed in my organization as a cloud, specifically as a SaaS, which is fully managed by Datafold. This means that hosting, maintenance, and automatic updates are all managed by Datafold, making it simpler for us and easier to get started.
The advice I would give others looking to use Datafold is that whoever is handling it, perhaps the head of IT, should have a sit-down with the analysts to ensure it fits into the stack that the organization is already conversant with.
Datafold is purpose-built for SQL and warehouse-based analytic pipelines with DBT, so if the current stack does not include a data warehouse and DBT, I would advise them to evaluate alternatives first. I also recommend using it for CI/CD quality, not just for general observability, because Datafold excels at pre-merge testing and data diffs.
Therefore, if the primary need is broad production or observability, the organization should also check out other options.
I rate Datafold a seven out of ten overall.
Information Technology and Services
Right one for testing
Reviewed on Apr 20, 2023
Review provided by G2
What do you like best about the product?
Awesome workflow, which is really a great feature of Datafold. Traditional way of doing data transfers is laborious. But this one helps to the best of capabilities. My fellow team was impressed.
What do you dislike about the product?
Nothing as such to say about negative here. But breadth of usage can be enhanced. This is not a drawback for sure. May be in coming days, will explore more and revert. For now all good
What problems is the product solving and how is that benefiting you?
Integration of data and managing data quality are two things that pose a big challenge in front of us. But Datafold made it easier to look and saved our time and effort
Telecommunications
Review for Datafold
Reviewed on Apr 19, 2023
Review provided by G2
What do you like best about the product?
Makes life easier with SQL code reviews helping find the hidden changes we made we dint know to our data.It gives ability to quality check the things in our own way with very much lesser errors compared to manual testing.
What do you dislike about the product?
Nothing in particular, a brief guide with documentation would have been justified. Helps in data managing of huge chunks of table, rows and records of data in day to day usage.
What problems is the product solving and how is that benefiting you?
Makes life way easier for automated testing Data accuracy and quality kpi achievement.easily integrates with modern data stacks such as Amazon redshift, snowflake,gitlab and GitHub