AWS Big Data Blog
Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 1: IAM-based access control
Your Google BigQuery users need to query data that lives in Amazon S3 Tables on AWS without copying it across clouds. This post shows how to connect BigQuery to Amazon S3 Tables through the AWS Glue Iceberg REST Catalog using IAM-based access control, so you keep one governed dataset and query it live from BigQuery.
Enable cross-cloud analytics with Amazon S3 Tables and Google BigQuery, Part 2: access control with Lake Formation
In Part 2 of this series, connect Google BigQuery to Amazon S3 Tables using AWS Lake Formation credential vending. Lake Formation manages fine-grained permissions and issues short-lived, scoped credentials to external engines, so you can centrally govern which teams and query engines read your Iceberg tables on AWS without managing IAM policies for every consumer.
GPU-accelerated Apache Spark with Amazon EMR and NVIDIA RTX PRO 4500 on Amazon EC2 G7 instances runs up to 3.7x faster
Amazon EMR on EKS now runs Apache Spark up to 3.7x faster on Amazon EC2 G7 instances with NVIDIA RTX PRO 4500 Blackwell GPUs than on comparable CPU instances, with no changes to existing Spark code. See the TPC-DS benchmark results, the cost comparison, and how to get started.
Introducing AWS Glue 6.0 for faster and more cost-effective data integration
AWS Glue 6.0 is now available, lowering AWS Glue pricing by 30%, adding an AWS optimized build of Apache Spark 4.1, and introducing Apache Iceberg V3 capabilities suitable for enterprise adoption. This post covers the key capabilities and performance benefits, with code examples to help you get started.
Upgrade AWS Glue jobs to Glue 6.0 with AI-powered Spark upgrades
Walk through upgrading a PySpark ETL job from AWS Glue 5.1 to AWS Glue 6.0 using the generative AI upgrades for Apache Spark. The upgrade analysis automatically detects incompatibilities, applies fixes, and validates results with data quality checks.
Long-term system tables retention in Amazon Redshift with Amazon S3 Tables
Amazon Redshift system table integration with Amazon S3 Tables automatically delivers your system table logs to Amazon S3 Tables in Apache Iceberg format. You can retain this data well beyond the 7-day limit for compliance, auditing, and cross-warehouse observability, without custom ETL pipelines or cluster resource consumption.
Track SageMaker Unified Studio project costs with custom tags and AWS CUR
Learn how to track Amazon SageMaker Unified Studio project costs by custom tags. This serverless solution enriches AWS Cost and Usage Report (CUR) data with custom project tags and visualizes cost by CostCenter, Team, or Environment in an Amazon Quick Sight dashboard.
Secure SageMaker Unified Studio access with SAML and conditional policies
Learn how to secure Amazon SageMaker Unified Studio by integrating it with an external SAML identity provider such as Okta. This post shows you how to apply conditional access policies that enforce device compliance, IP-based restrictions, and multi-factor authentication for your data and AI workloads.
Querying raw log data using SQL and PPL with the optimized engine in Amazon OpenSearch Service
Learn how to run fast analytical queries directly against raw log and trace data in Amazon OpenSearch Service using PPL and SQL. Follow a single incident investigation, one query at a time, and see how the new optimized engine answers each question directly from raw spans.
Fresher insights, faster decisions: talabat’s near-real-time analytics across AWS and Google Cloud
Leading everyday app across the Middle East and North Africa, talabat, built a hybrid multi-cloud lakehouse that keeps a single Apache Iceberg copy of streaming data on Amazon S3 Tables while letting Google BigQuery query it in place, eliminating cross-cloud data duplication and schema-synchronization overhead.









