
Sold by: MERMAID
Open data
|
Deployed on AWS
Community-sourced repository of coral reef image classification training data, including continually updated confirmed annotations from [MERMAID](https://datamermaid.org/)
Overview
Community-sourced repository of coral reef image classification training data, including continually updated confirmed annotations from MERMAID
Features and programs
Open Data Sponsorship Program
This dataset is part of the Open Data Sponsorship Program, an AWS program that covers the cost of storage for publicly available high-value cloud-optimized datasets.
Pricing
This is a publicly available data set. No subscription is required.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Legal
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
Delivery details
AWS Data Exchange (ADX)
AWS Data Exchange is a service that helps AWS easily share and manage data entitlements from other organizations at scale.
Open data resources
Available with or without an AWS account.
- How to use
- To access these resources, reference the Amazon Resource Name (ARN) using the AWS Command Line Interface (CLI). Learn more
- Description
- The coral-reef-training AWS S3 bucket provides a single, open, well-structured, growing, community-sourced repository of coral reef image classification training data. Hosted at s3://coral-reef-training, this bucket supports global efforts in coral reef conservation through standardized, machine-learning-ready imagery and annotations. The bucket serves as the image storage backend for MERMAID’s image classification workflows and to distribute confirmed and scrubbed MERMAID coral reef image data, but it also provides a shared location where partners including CoralNet can contribute to and benefit from collective ML model development, each according to its own data structures and policies. Data in the bucket is free and open for public access; only contributing organizations have write access to their own data prefixes. By centralizing and standardizing coral reef image data, this initiative accelerates collaboration across scientific, conservation, and machine learning communities and facilitates the creation of a common, evolving image classification model for coral reefs worldwide.
- Resource type
- S3 bucket
- Amazon Resource Name (ARN)
- arn:aws:s3:::coral-reef-training
- AWS region
- us-east-1
- AWS CLI access (No AWS account required)
- aws s3 ls --no-sign-request s3://coral-reef-training/
Resources
Vendor resources
Support
Contact
Managed By
How to cite
Community coral reef image classification training data was accessed on DATE from https://registry.opendata.aws/coralreef-image-classification-training .
Similar products
Training AI models involved in industrial use cases requires a large amount of data/images, which can be very long or even impossible to acquire in the field.
Indeed, highly reliable systems will rarely generate data associated with failures, which our customers seek to prevent.
Generating this data (especially images) can be tedious, even with image/data generation tools, because these tools must be configured/parameterized to generate a large amount of data with particular attention to borderline cases.
To effectively meet these business needs, our teams are implementing and continuing to develop for our customers a methodology and assets that takes advantage of LLMs, image and data generation tools, quality measurement and correction algorithms co-developed as part of Confiance.ai as well as several AWS services.
Pre-configured Amazon Machine Image with PyTorch 2.1 and CUDA 12.1 for accelerated deep learning. This production-ready environment eliminates complex setup processes, saving 4+ hours of configuration time. Includes full GPU optimization for NVIDIA hardware, essential ML libraries, and security configurations out-of-the-box. This product wherein additional charges apply for support provided by Galaxys.
Ideal for researchers, data scientists, and developers working on computer vision, natural language processing, and neural network projects. Features automatic environment setup, Jupyter Lab integration, and optimized performance for AWS EC2 instances.
Classify Document AI is a production-ready AI endpoint that identifies and categorizes image documents (PDFs, scans, or images) by type so each one is routed to the right workflow automatically. Output is clean, consistent, and ready for automation.
Centific AI Datasets delivers production-ready, enterprise-grade training data for AI and machine learning models. Powered by 1.8 million domain experts across 230+ languages, our custom datasets span text, speech, image, video, and synthetic modalities to accelerate your AI development.