Customer work · Financial services

DHFL Pramerica used an AWS data lake to turn policy data into customer risk profiles

DHFL Pramerica needed a consolidated customer view and a repeatable way to assess risk before recommending insurance products. BluePi built a machine-learning path from life-policy data through AWS Glue, Amazon S3, Amazon EMR, and Amazon Redshift. The system produced high-, medium-, and low-risk customer segments from demographic, education, health, and policy attributes.

DHFL Pramerica used an AWS data lake to turn policy data into customer risk profiles system diagram

Customer risk profiling for life insurance decisions

The insurer needed customer attributes from life-policy data to support product-selection and underwriting workflows. BluePi connected on-premises DB2 data to AWS through VPN, loaded raw data into Amazon S3 with AWS Glue, trained the profiling model with Amazon EMR, and persisted processed data and model artifacts in Amazon S3.

01 · DB2-to-AWS data movement

On-premises DB2 life-policy data reached AWS through a VPN, with AWS Glue loading raw data into Amazon S3.

02 · Risk model attributes

The model used demographic, education, health, and policy attributes to classify customers as high, medium, or low risk.

03 · Training and serving split

Amazon EMR trained the model and served the resulting profiles, while Amazon Redshift supported analytical use.

A controlled AWS path from policy data to risk profile

DB2 life-policy data moved to AWS through VPN. AWS Glue loaded raw records into Amazon S3, Amazon EMR trained the profiling model, Amazon S3 stored processed data and model artifacts, Amazon Redshift supported analysis, and Amazon EMR served profiles to downstream use.

Separate raw and processed data

Distinct Amazon S3 layers preserved the source boundary while giving training and serving paths controlled model inputs.

Train and serve through distinct paths

Amazon EMR training produced the model artifacts, while a separate Amazon EMR serving path supplied profiles to downstream use.

Keep planned sources explicit

DB2, CSV, Adobe Experience Manager, and LinkedIn inputs are scoped for phase two; the delivered system runs on the life-policy path.

Where this pattern fits

This case fits insurers that need a repeatable path from policy data to a customer risk profile used in product selection or underwriting. BluePi can begin with one decision, the attributes it requires, and the system boundary that must deliver the profile.

Discuss this case with BluePi

Case details

Open a section to review the customer problem, implementation, business change, and architecture.

01The starting point

Customer information could not be combined into a complete view for risk assessment. DHFL Pramerica needed to calculate customer risk aversion consistently and connect the resulting profile to relevant insurance products.

  • Fragmented customer evidence: Life-policy and related customer attributes were difficult to assess together.
  • Manual risk assessment: Underwriting lacked an automated customer risk-profile input.
  • Generic product selection: Teams could not consistently match insurance products to an individual risk profile.
02System delivered

BluePi connected on-premises life-policy data to AWS through a VPN. AWS Glue loaded raw data into Amazon S3, Amazon EMR trained the profiling model, and the data-lake path persisted processed data and model artifacts. Amazon Redshift supported analytical use, while Amazon EMR served the resulting profiles.

  • Central data-lake path: AWS Glue extracted source data into an Amazon S3 raw zone, while a separate Amazon S3 data-lake layer held processed data and model artifacts.
  • Machine-learning workflow: Amazon EMR trained profiles from demographic, education, health, and policy attributes. Amazon Redshift supported analysis, and Amazon EMR served the model output.
  • Operational risk segments: The system classified customers as high, medium, or low risk for product-selection and underwriting use.
03Outcomes

DHFL Pramerica gained a reusable customer risk profile for product-selection and underwriting workflows. Risk segments made the basis for product recommendations explicit, while automated profiling reduced the manual work required before underwriting decisions.

  • Complete customer risk view: The consolidated data and model output created one reusable risk profile for each customer.
  • Risk-informed product selection: Teams could use the risk score and segment when selecting a relevant insurance product.
  • Automated underwriting input: Risk profiling became an automated input to underwriting, reducing the manual assessment work required before an application could progress.
04Architecture boundary

On-premises DB2 life-policy data crossed a VPN into AWS Glue extraction and an Amazon S3 raw-data zone. Amazon EMR trained the model, while Amazon S3 persisted processed data and model artifacts. Amazon Redshift supported analytical use, and a separate Amazon EMR path served the resulting profiles.

05Planned source expansion

Additional DB2 records, CSV files, Adobe Experience Manager, and LinkedIn inputs are scoped for phase two. Those planned sources stay outside the delivered life-policy path, so scope stays unambiguous.

06What changed for product and underwriting teams

Product and underwriting teams gained a shared customer risk profile instead of assembling an assessment from disconnected attributes. The high-, medium-, and low-risk segments connected model output to the insurance-product and underwriting workflows that used it.

System diagram