Overview

Datralys Inc. is a US-based enterprise analytics firm managing large-scale, multi-source data ecosystems for Fortune-segment clients across finance, logistics, and healthcare verticals. When Datralys engaged Bacancy, their ingestion pipelines were brittle standalone scripts with no orchestration or monitoring, data assets had no metadata governance layer, and search was entirely keyword-based, returning inconsistent results against large heterogeneous datasets. Bacancy modernized the entire data infrastructure by building a production-grade orchestrated pipeline using Apache Airflow and dbt, implementing centralized metadata governance via Apache Atlas, integrating cloud data ingestion through AWS Glue, and introducing an OpenAI Embeddings layer indexed against Elasticsearch to deliver intelligent, high-precision data retrieval across the Datralys platform.

Technologies Used

Python
Apache Airflow
dbt
Elasticsearch
AWS Glue
OpenAI Embeddings

Project Highlights

checkmark

Automated end-to-end data pipeline with Airflow and dbt

checkmark

AI-powered metadata layer for semantic search

checkmark

Unified metadata governance via Apache Atlas integration

checkmark

40% retrieval precision lift with vector and full-text search

The Challenges

1

Fragmented data sources with inconsistent schemas made unified ingestion and normalization extremely complex. The client's data was spread across multiple systems and platforms, each using different formats and structures, making integration difficult and time-consuming.

2

Absence of a metadata governance framework led to poor data discoverability and unreliable retrieval outcomes. Teams struggled to locate trusted datasets and lacked visibility into data lineage, ownership, and quality.

3

Legacy pipeline architecture caused high latency and frequent failures during peak data ingestion windows. The client's existing workflows were difficult to scale and lacked the reliability needed to support growing data volumes.

4

No semantic search capability meant keyword-based queries consistently returned irrelevant or incomplete data results. Users had to rely on exact keywords and dataset names, making it difficult to find relevant information across the data ecosystem.

Solutions by Bacancy

1

We implemented a schema-agnostic ingestion layer using AWS Glue and Python that abstracted source-specific differences behind a unified normalization interface. We designed a configuration-driven framework where each data source registers its schema contract separately from the core pipeline logic. Schema drift from upstream sources now triggers validation alerts rather than silent failures, and new data sources can be onboarded without touching the core pipeline.

2

Our team deployed Apache Atlas as the centralized metadata catalog. We configured it to automatically capture lineage events at every pipeline stage, from raw ingestion through dbt transformations to the final serving layer. Automated tagging policies classify datasets by domain, sensitivity, and quality score without manual intervention, giving every analyst full provenance visibility before they touch a dataset.

3

Bacancy re-architected the pipeline using Apache Airflow, replacing flat cron scripts with a fully dependency-aware DAG system. Every task now declares its upstream dependencies, Airflow enforces execution order, and failures halt downstream work automatically. We added retry logic with exponential backoff, SLA-based alerting, and layered dbt on top for modular, testable SQL transformations with built-in data quality assertions at every run.

4

Our data scientists integrated OpenAI Embeddings with Elasticsearch to replace keyword search with a semantic retrieval layer that understands query intent. Every dataset along with its Atlas metadata, glossary definitions, and field descriptions was embedded into a vector space and indexed in Elasticsearch. The hybrid retrieval approach combines semantic similarity with keyword relevance, surfacing the right datasets even when query terms do not literally match field names or dataset titles.

Core Features

checkmark

AI-Powered Semantic Search Engine for Data Discovery

checkmark

Automated Metadata Governance and Lineage Tracking

checkmark

Resilient Pipeline Orchestration with Real-Time Alerting

checkmark

Scalable Cloud Ingestion Architecture for Enterprise Growth

No. of Resource

04

No. of Resource

Time Frame

July 2025 - February 2026

Time Frame

Project Snapshot

How We Built an AI-Ready Data Pipeline and Metadata Layer That Lifted Retrieval Precision by 40%

Outcomes

40% improvement in data retrieval precision across pipelines

55% reduction in end-to-end pipeline latency post optimization

100% of critical metadata assets cataloged and governed centrally

60% faster onboarding for new data consumers and teams

Automated lineage tracking eliminated all manual data audit efforts

3x scalability achieved on AWS for growing enterprise data volumes

Technical Stack

DATA PIPELINE ORCHESTRATION Apache Airflow
DATA TRANSFORMATION dbt (Data Build Tool)
SEARCH & RETRIEVAL ENGINE Elasticsearch
CLOUD DATA INTEGRATION AWS Glue
AI / EMBEDDING LAYER OpenAI Embeddings API
SCRIPTING & AUTOMATION Python 3.10
DATA WAREHOUSE Amazon Redshift
METADATA MANAGEMENT Apache Atlas
VERSION CONTROL & CI/CD GitHub Actions
PROJECT & ISSUE TRACKING Jira

Experience With Bacancy

2500+ Projects Experienced Innovation with Bacancy!

Get access to an experienced team of developers and engineers from Bacancy, handpicked to ace your goals. Kickstart within 48 hours, no-risk trial.

Book a 30 min call

14+

Years of Business Experience

1458+

Happy Customers

12+

Countries with Happy Customers

1050+

Agile Enabled Employees

How Can We Help?