The PDF Obfuscation Pipeline is a powerful solution for transforming sensitive PDF documents into safe, shareable assets. It enables organizations to unlock the value of clinical data while ensuring strict compliance with HIPAA, GDPR, and institutional privacy standards.
The PDF Obfuscation Pipeline helps healthcare providers, researchers, and data scientists safely access and use sensitive clinical data without compromising patient privacy by masking PHI information in PDFs.
Whether you are training machine learning models, building decision-support tools, or enabling research collaborations, our PDF Obfuscation Pipeline ensures your documents are privacy-compliant, legible, and faithful to their original structure.
Masked entities include HOSPITAL, NAME, PATIENT, ID,MEDICALRECORD, IDNUM, COUNTRY, LOCATION, STREET, STATE, ZIP, CONTACT, PHONE, DATE. The output is a PDF document, similar to the one at the input, but with fake obfuscated text on top of the targeted entities.
IMPORTANT USAGE INFORMATION:
After subscribing to this product and creating a SageMaker endpoint, billing occurs on an HOURLY BASIS for as long as the endpoint is running.
-Charges apply even if the endpoint is idle and not actively processing requests.
-To stop charges, you MUST DELETE the endpoint in your SageMaker console.
-Simply stopping requests will NOT stop billing.
This ensures you are only billed for the time you actively use the service.
Highlights
**Entity-Level Obfuscation**
The pipeline performs targeted obfuscation of sensitive entities such as NAME, PHONE, and more. These entities are replaced by realistic surrogates, preserving the document usability while ensuring no original data leaks.
**Customizable Entity Scope**
You define what matters. The pipeline allows full customization of which entities to obfuscate or preserve, giving you control over your de-identification strategy.
**Rendering-Aware Replacement**
Replacements are layout-aware. A name John Smith will be replaced with similar-length entity like Mike Burke, ensuring the document remains visually consistent and readable.
**Consistent Entity Replacement**
If John Smith is replaced by Mike Tyson on page 1, all other instances of John Smith on all pages will be replaced consistently, preserving referential integrity, which is critical for longitudinal or document-linked analysis.
**Date Shifting Support**
You can apply coherent temporal transformations, such as shifting all dates by 2 months, while preserving the internal temporal relationships between events.
**Open Evaluation Dataset**
We built and released a benchmark dataset to help evaluate document-level de-identification. It includes metrics and examples that showcase what this pipeline can achieve in realistic clinical settings.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by the hour based on the compute instance you run and the processing mode you choose. Three instance sizes are available: ml.m4.4xlarge, ml.c5.4xlarge, and ml.c5.9xlarge. Larger instances offer more compute for heavier workloads. Each instance comes in two modes. Batch mode processes documents in scheduled groups. Real-time mode handles requests as they arrive. This gives you six combinations of instance size and mode. Your total cost depends on which instance you pick, which mode you use, and how many hours you run it.
Top-of-mind questions for buyers
What does one billed hour cover, and am I charged when the instance is stopped?
You pay per host-hour for the instance you run. Charges accrue only while the instance is running and processing. A stopped or terminated instance stops software metering. Underlying AWS resource fees may still apply based on your AWS account setup.
How do batch mode and real-time mode differ for my bill?
Both meter per running host-hour on the same instance. Batch mode processes documents in scheduled groups, so you can start it, run a job, and stop it. Real-time mode stays available to handle requests as they arrive, so hours accrue while the endpoint runs. Batch suits periodic jobs; real-time suits continuous demand.
How do I control cost as my document volume grows?
Cost scales with running hours, not document count directly. You raise throughput by choosing a larger instance or running longer. The ml.c5.9xlarge offers more compute than the ml.c5.4xlarge or ml.m4.4xlarge for heavier workloads. You can also run batch jobs and stop the instance between runs to limit accrued hours.
www.johnsnowlabs.com
Helpful?
Vendor refund policy
No refunds are possible.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An Amazon SageMaker model package is a pre-trained machine learning model ready to use without additional training. Use the model package to create a model on Amazon SageMaker for real-time inference or batch processing. Amazon SageMaker is a fully managed platform for building, training, and deploying machine learning models at scale.
Deploy the model on Amazon SageMaker AI using the following options:
Real-time inference
Deploy the model as an API endpoint for your applications. When you send data to the endpoint, SageMaker processes it and returns results by API response. The endpoint runs continuously until you delete it. You're billed for software and SageMaker infrastructure costs while the endpoint runs. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Deploy models for real-time inference .
Batch transform
Deploy the model to process batches of data stored in Amazon Simple Storage Service (Amazon S3). SageMaker runs the job, processes your data, and returns results to Amazon S3. When complete, SageMaker stops the model. You're billed for software and SageMaker infrastructure costs only during the batch job. Duration depends on your model, instance type, and dataset size. AWS Marketplace models don't support Amazon SageMaker Asynchronous Inference. For more information, see Batch transform for inference with Amazon SageMaker AI .
Version release notes
Upgrade underlying JSL libraries to version 6.4.0
Additional details
Inputs
Outputs
Usage instructions
Sample notebooks
Inputs
Summary
Supported Content Types:
application/octet-stream
application/pdf
image/png
image/jpeg
image/bmp
image/tiff
image/gif
Notes
Input must be a valid PDF or image file sent as raw bytes.
Both single-page and multi-page PDFs are supported.
application/octet-stream is accepted for any supported format — the model detects the file type automatically from the binary content.
For batch transform jobs, upload each file as a separate S3 object and use SplitType: None with BatchStrategy: SingleRecord to prevent binary splitting.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.