The ESG Book Estimated Emissions Data Module provides investors with estimated emissions for ~45,000 public corporate entities that do not disclose their emissions. The dataset includes estimations for Scope 1, Scope 2, Scope 3 (total) emissions, and 15 Scope 3 Categories. A confidence rating is also provided alongside each estimated emissions figure, indicating the degree of accuracy of the estimation based on the amount of available data used in the estimation process.
The ESG Book Estimated Emissions Data Module provides investors with estimated emissions for ~45,000 public corporate entities that do not disclose their emissions. The dataset includes estimations for Scope 1, Scope 2, and Scope 3 (total) emissions, as well as the 15 Scope 3 Categories in tonnes of CO2 equivalents. A confidence rating is also provided alongside each estimated emissions figure, indicating the degree of accuracy of the estimation based on the amount of available data used in the estimation process. Importantly, PCAF data quality indicators are included.
The dataset additionally includes the actual reported emissions data of public companies. We currently cover 4000 public companies, where approximately half of them disclose their emissions data. Our Estimated Emissions Data Module thus significantly expanding the coverage of emissions data for use in portfolio analysis and index creation, for instance.
Methodology Overview
Emissions are estimated using the Extreme Gradient Boosting (XGBoost) Model. The model is an unsupervised machine learning model which identifies and analyses complex relationships between large numbers of predictor variables to generate estimations for unknown data. In this case, the model identifies the relationship between 15 financial and non-financial predictor variables and emissions for each region, country, sector and industry to estimate the emissions of companies which are not disclosing emissions data.
We have chosen to use a machine learning estimation model rather than a traditional statistical regression model for several key reasons. Firstly, the XGBoost model (machine learning model) is able to handle non-linear relationships. As the predictor variables might be non-linearly correlated with emissions (for instance, a company with 500 employees might not generate 5 times the emissions of a company with only 100 employees due to economies of scale), the ability of the XGBoost model to handle non-linear relationships provide an extra layer of robustness to accurately capture the relationships between the predictor variables and emissions.
Secondly, the XGBoost model is able to handle missing data unlike conventional regression models or other machine learning models such as Adaptive Boosting. Though 15 predictor variables are used in the model, all 15 datapoints might not be available for all companies. As such, a threshold of datapoints is set such that the model will estimate emissions for companies which meet this minimum data threshold. Conventional regression models are unable to account for this missing data, where this missing data has to be interpolated, or simply replaced with zeros. This introduces higher order errors into the model, reducing the accuracy of the emissions estimations due to the ambiguity of input data. This issue does not affect the XGBoost model due to its ability to handle missing data.
Lastly, the XGBoost model uses a decision-tree algorithm to identify and analyse the complex relationships between the predictor variables and emissions, which is subsequently used in the estimation process. This allows for greater accuracy as the decision tree process corrects the mistakes of the previous trees. The parameters of the model are fine-tuned to increase the precision of estimations. This is done using the Optuna4 , an open source hyperparameter optimization framework, that tests different configurations of hyperparameters on a holdout test set to determine the optimal values for a given regression.
Overall, due to the reasons explained above, the XGBoost model shows better accuracy when compared to traditional statistical models such as the Ridge Regression model or other machine learning models such as the Adaptive Boost model.
Use Cases
The Emissions data can be instrumental for Asset Managers and Corporates:
Portfolio Management
Emissions data can be used by Portfolio Managers during portfolio construction for:
Exclusion - The screening out of companies that are not aligned with the Paris Agreement temperature goals.
Carbon Intensity - The scaling of emissions data by financial metrics to compute carbon intensities, monitor the portfolio and benchmark against other portfolios
TCFD & SFDR Reporting - The reporting of climate-related financial metrics to understand the climate-related risks and opportunities of the companies within a portfolio
Portfolio Alignment to Climate Goals - Identify to what extent a portfolio is aligned with the Paris Agreement to minimise exposure to carbon-intensive companies
Regulation Compliance - Generate voluntary TCFD disclosures on how climate-related risks and opportunities are factored into relevant investment strategies
Alignment to investor demand- Increasing number of investors require asset managers to integrate climate risks and opportunities into their investing strategy
Corporates
Emissions data includes climate metrics related to emissions, reporting, policy and frameworks and enables:
Tailored benchmarking - The quality and granularity of the data allows corporations to analyse their climate performance against direct peers, industry, sector and region
Climate Reporting - The identification of climate-related topics that need to be reported on for a company to stay ahead of its peers
Tailor-made Comparison Metrics - Combining emissions data with financial metrics such as revenue or EBITDA or non-financial metrics such as production quantity enables the creation of innovative carbon intensity metrics relevant to each company
Market Positioning & Differentiation - Understand which climate-related topics corporations need to report on to be a leader among their peers
Metadata
Meta Data
Information
Update Frequency
Weekly
Data Source(s)
Estimations produced using ESG Book raw emissions data and ESG raw data. Financial data from third-party provider
Geographic coverage
Global
Time period coverage
Present
Is historical data “point-in-time”
YES
Raw or scraped data
All input data is collected from public sources such as Annual Reports, CSR Reports, Investor Relation Presentations and Reports, ESG reports, Company Websites
Number of companies covered
~37,000
Standard entity identifiers
Ticker (please contact for information on other identifiers)
Pricing Information
Pricing is determined on a use-case basis, thus please contact for more information.
When requesting please include the following information:
Organization Name
Position (non-mandatory)
Business Email Address or Telephone Number
Country
Use-case
Regulatory and Compliance Information
This product is allowed for internal use only, users are not allowed to distribute the data externally.
If you're interested in a re-distribution of data use case, please contact us.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on the duration and terms of your contract with the vendor. This entitles you to a specified quantity of use for the contract duration. If you choose not to renew or replace your contract before it ends, access to these entitlements will expire.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
This listing uses a single pricing dimension called Product Access (Units). You buy access through a contract, and the units grant subscribers entry to the product. There are no separate tiers, instance sizes, or usage-based add-ons to choose between. Pricing scales with the number of access units you commit to under the contract. The product delivers validated climate and sustainability data, including Scope 1, 2, and 3 emissions estimates, which you can integrate through an API, common data-sharing platforms, secure file transfer, CSV files, or a web platform.
Top-of-mind questions for buyers
What does one unit of Product Access represent for billing?
A unit grants a subscriber access to the emissions data module under your contract. Access units map to subscriber entry to the product, not to a fixed data volume or company count. You commit to a number of units, and pricing scales with that quantity.
How can I access the emissions data once I subscribe?
You can integrate the data through an API or common data-sharing and warehouse platforms. You can also use secure file transfer, download CSV files, or view data directly on the vendor's web platform. Public data refreshes daily, and private data updates as new information arrives.
What data scope is covered under the Product Access contract?
The module provides validated climate and sustainability data, including Scope 1, 2, and 3 emissions estimates. Data covers company-level metrics with historical records, unit standardisation, and source documentation for each data point. For details on which datasets your units include, contact the vendor.
www.esgbook.com+1
Helpful?
Vendor refund policy
No refunds.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
The Decarbonisation Analytics assesses 45,000+ companies comprehensively on reporting quality, climate impact
and target alignment using the latest IEA scenarios benchmark data.
Our SFDR Principal Adverse Impact (PAI) Indicators Solution has been designed by ESG Book to enable clients to meet the reporting requirements of the EU Sustainable Finance Disclosure Regulation (SFDR), which entered into force on 10 March 2021. ESG Book’s SFDR PAI Solution is a granular tool that empowers clients to ingest the company-level data indicators needed for SFDR reporting based on our in-house SFDR raw data calculation methodology.
The Risk Score provides investors and corporates with a transparent and systematic assessment of the
sustainability exposure of corporate entities. The Risk Score measures company exposures, using a normative
assessment, relative to universal principles of corporate conduct as defined by the ten principles of the UN’s
Global Compact (GC).