Data is not perfect. Refining data without comprehensive and exact data insights is costly. Our platform reduces your efforts, time and cost significantly by letting data tell its own story: full fidelity metadata insights.
Full fidelity metadata sans sampling generates a blueprint of source data, comprehensive, deep and exact insights, for AI, ML, data science, analytics, governance, test data, and more. Brings a unique combination of automated discovery and analysis with a processing engine in one solution to control escalating compute costs. Metadata discovery, analysis and quality metrics are through a no code, drill down GUI, nLite. If you wish to develop custom solutions using our products we offer f2md API Toolkit for programming.
The product is data store agnostic and hybrid, working directly on files in object stores, data lakes or databases, without data movement. Its AI driven metadata engine f2mdbX helps you discover schema and quality anomalies even on unknown, legacy datasets resident in cloud. It learns your source data in production pipelines to identify and inform you of changes in schema, data types, values or statistics for proactive data monitoring. Metadata on source data extends beyond profiling to rules, referential integrity, accuracy and privacy to identify sensitive data values.
Available as a self-hosted solution to run on single or multi-node clusters, product is scaleable to suit your data needs and is optimized to work well on a few large files from data warehouses, or thousands of small files from IoT.
Highlights
Discovery: Hyper-efficient processing of every value in a dataset to generate comprehensive and deep metadata. A bottom-up approach to let data tell its own story without bias or ambiguity.
Analysis: Insights derived from full fidelity metadata to identify data issues that are not easily found by sample based metadata or by running a plethora of SQL queries on source data.
Performance: Runs as a distributed metadata engine on AWS clusters, taking advantage of both compute and data parallelism, to deliver high performance for both discovery and analysis.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour based on the AWS instance size you run this SQL data engine on. The nine options fall into two instance families. The m6i and m5 sizes give you more memory relative to compute. The c6i sizes give you more compute relative to memory. Within each family, larger sizes carry more cores and memory, so the hourly rate rises accordingly. You can scale up to a bigger instance or scale out by adding nodes to match your workload. Billing runs only while an instance is active.
Top-of-mind questions for buyers
What resources determine the hourly rate for each instance option?
The rate follows the AWS instance you launch. The m6i and m5 options give more memory relative to compute. The c6i options give more compute relative to memory. Larger sizes within a family carry more CPU cores and memory, so their hourly rate is higher.
Am I charged when an instance is stopped or paused?
Hourly software charges accrue only while an instance runs. Fully stopped instances stop the software meter. You may still pay underlying AWS storage fees for stopped instances, but the software licence meters running time only.
How do I add capacity as my data workload grows?
You can scale up by moving to a larger instance size, or scale out by adding nodes. Compute and storage scale independently. Systems can grow one node at a time, so you adjust capacity to match your workload as it changes.
www.xtremedata.com
Helpful?
Vendor refund policy
We do not currently support refunds, but you can cancel at any time.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Hyper-efficient processing of every value in a dataset without sampling to generate comprehensive and deep metadata insights across all data points.
Data Store Agnostic Processing
Works directly on files in object stores, data lakes, or databases without requiring data movement, supporting hybrid deployments across multiple storage types.
AI-Driven Schema and Anomaly Detection
Automated discovery engine that identifies schema changes, data type variations, value anomalies, and statistical deviations in unknown or legacy datasets through machine learning.
Distributed Metadata Engine Architecture
Runs as a distributed processing engine on AWS clusters leveraging compute and data parallelism to deliver high-performance metadata discovery and analysis at scale.
Comprehensive Data Quality Assessment
Evaluates data quality through rules validation, referential integrity checks, accuracy metrics, and sensitive data identification beyond standard profiling capabilities.
Universal Query Engine
AutoSQL provides a universal query engine for unified data access across disparate data sources.
Data Discovery and Classification
Watson Knowledge Catalog enables real-time discovery and classification of data with automated cataloging capabilities.
Automated Policy Enforcement
Pervasive privacy framework with automated policy enforcement for sensitive data protection across all users in the organization.
Model Operations and Governance
ModelOps on Watson Studio synchronizes application and model pipelines while monitoring and governing AI models to manage risk, reduce drift and bias, and enhance transparency.
Data Fabric Architecture
Data fabric technology connects siloed data on premises or across multiple clouds without requiring data movement, enabling consolidated and governed views of enterprise data.
Metadata Centralization
Centralizes metadata from disparate sources into a unified platform for discovering, describing, governing, and managing data assets including data, BI reports, and AI models.
Behavioral Analysis Engine
Incorporates a Behavioral Analysis Engine to provide advanced analytics and insights across data assets.
Data Lineage and Tracking
Enables documentation of insights and tracking of data lineage across teams for transparency and compliance purposes.
Self-Service Analytics
Supports self-service analytics capabilities allowing users to independently discover and analyze data assets.
AI Governance Framework
Provides an AI governance framework that ensures data quality, transparency, and compliance for AI initiatives.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.