← Back to Dashboards

Case Study

CFPB Consumer Complaint Analysis

How a cloud-connected, auto-refreshing pipeline turns 8.3M+ raw complaint records into a live executive dashboard — with zero manual data handling.

DatabricksDatabricks Power BIPower BI Automated PipelineScheduled API Refresh

The Problem

The Consumer Financial Protection Bureau publishes millions of consumer complaint records, but a raw dataset of this size — over 8.3 million rows spanning January 2025 through July 2026 — isn't something a business leader can read directly. Without a reporting layer, systemic issues like which companies generate the most complaints, or which regions see the slowest response times, stay buried in the data.

8.3M+
Complaints Analyzed
7.1M+
Tied to the "Big Three" Bureaus
4.3M+
Linked to Incorrect Info

How the Pipeline Works

The defining feature of this project is that it's not a static report — it's a live system. Data flows from source to dashboard automatically, on a schedule, without anyone touching a spreadsheet.

1

Scheduled Ingestion

A scheduled job calls the Databricks API to pull the latest complaint records directly from the cloud workspace — no manual exports or file transfers.

2

Transform & Clean

Records are deduplicated, standardized, and structured into a reporting-ready model, isolating the fields that matter most for trend and root-cause analysis.

3

Load

Cleaned data lands in a reporting layer optimized for fast queries, so the dashboard stays responsive even at millions of rows.

4

Publish to Power BI

Power BI connects directly to the refreshed dataset. Each cycle, the dashboard updates itself — leadership always sees current numbers, not a stale snapshot.

What This Demonstrates

Key Findings

View the Live Dashboard