Skip to content
Writing
Esc
Type to search all articles.

    Debanjan Saha

    Senior Machine Learning Engineer & Data Scientist

    Summary

    Seasoned Machine Learning Engineer with 10+ years of experience architecting Generative AI, LLM, and multimodal ML systems across multi-cloud environments (AWS, Azure, GCP). Adept at end-to-end ML lifecycle management, model optimization, and real-time AI deployment. Proven leader in cross-domain projects spanning marketing analytics, financial risk modeling, and healthcare AI. Seeking Lead Machine Learning Engineer role advancing foundation model development and scalable AI infrastructure.

    10+

    Years ML & AI Experience

    ~45 TB/day

    Data Pipelines Scale (Mastercard)

    10M+

    Documents Digitized via Multimodal AI (AT&T)

    4.0

    GPA, MS Data Science (Northeastern)

    Skills

    Spotlight roles by:

    ML Frameworks
    PyTorch, TensorFlow, Keras, Scikit-learn, Spark ML, Hugging Face, OpenAI, LLM, RAG, LangChain, LlamaIndex, MCP
    DevOps & MLOps
    Docker, Kubernetes, AKS, EKS, GKE, Helm, Terraform, MLflow, Kubeflow, ELK, Prometheus, Databricks
    Cloud Ecosystems
    AWS (SageMaker, S3, EMR), Microsoft Azure (Azure ML, ADLS, AI Search), GCP (Vertex AI, Cloud Run)
    Programming & Databases
    Python, Scala, PySpark, Matlab, C/C++, SQL (PostgreSQL, MS-SQL, MySQL), HTML/CSS, React, Node.js
    Data Visualization & Analytics
    Tableau, PowerBI, QuickSight, Grafana, Looker, Jupyter Notebook, Data Storytelling

    Experience

    1. Mastercard

      Senior Data Engineer

      January 2025 - Present

      Dublin, Ireland

      • Engineered large-scale data pipelines processing ~45 TB of daily transaction data across Mastercard systems, enabling advanced analytics use cases including segmentation, lifetime value, churn, affinity, loyalty, and share-of-wallet modeling.
      • Contributed to the development of Pioneer, an internal privacy-enhanced analytics platform powering 200+ Mastercard users, providing real-time access to curated datasets with a Zeppelin-based notebook interface.
      • Architected secure, high-performance data access layers replacing legacy Hadoop workflows, reducing data retrieval latency by >70% and improving query reliability.
      • Prototyped Text-to-SQL capabilities into Pioneer using LLMs, enabling natural language query generation to democratize data access.
      • Designed and maintained ML pipelines in Azure Databricks and MLflow, supporting end-to-end model lifecycle management.
      • Ensured compliance with PCI-DSS and GDPR standards while enabling secure feature engineering and model deployment at scale.
      • Advocated internal ML infrastructure evolution, driving automation using Terraform, reproducibility, and cost optimization via Photon compute, Delta Lake optimization, and parallelized data ingestion.
      • PySpark
      • Azure Databricks
      • Delta Lake
      • MLflow
      • Terraform
      • LLM
      • Text-to-SQL
      • PCI-DSS
      • GDPR
    2. AT&T

      Machine Learning Engineer

      September 2024 - January 2025

      Dallas, TX, USA

      • Built an end-to-end multimodal AI pipeline to digitize over 10M+ legacy tower contract PDFs, leveraging Vision Transformers, PyMuPDF, and OCR models (Tesseract, OCRmyPDF), improving accuracy by 34% over baseline OCR.
      • Benchmarked and integrated state-of-the-art embedding models (OpenAI, Azure, Sentence-Transformers), increasing retrieval precision by 42% in Azure AI Search.
      • Architected an Azure-based knowledge retrieval system storing embeddings and document metadata in AI Search + ADLS, enabling instant semantic search and metadata-aware insights.
      • Fine-tuned a LLaMA-based Large Language Model on internal contract embeddings and metadata, reducing legal drafting time by ~60%.
      • Implemented query tokenization, reranking, and Retrieval-Augmented Generation (RAG) in production search UI, improving relevance and latency by >25%.
      • Deployed scalable inference endpoints on Azure Kubernetes Service (AKS) and integrated APIs into the ContractsAI web interface.
      • LLaMA
      • RAG
      • Vision Transformers
      • Azure AI Search
      • Azure Kubernetes Service (AKS)
      • PySpark
      • Azure Databricks
      • OpenAI Embeddings
    3. Nanobiosym

      Principal AI Engineer

      May 2023 - December 2023

      Cambridge, MA, USA

      • Architected a cloud-agnostic Generative AI platform on Kubernetes (EKS) and Docker within a HIPAA-compliant AWS environment, delivered in under 6 months for multi-tenant scalability.
      • Developed LSTM-based COVID-19 severity prediction models leveraging patient vitals and longitudinal data, reducing intervention response time by 30%.
      • Fine-tuned domain-specific NER models on patient–clinician dialogues, achieving an F1 score of 0.87 in symptom extraction and severity assessment.
      • Deployed a conversational AI assistant using Amazon Lex and Polly, enhancing patient engagement by 40% through real-time symptom triage.
      • Implemented shadow binning and SNS-triggered alerting for critical patient classification, cutting emergency response times by 20%.
      • AWS (EKS, SageMaker, Lex, Polly, SNS)
      • Generative AI
      • LSTM
      • NER
      • Docker
      • HIPAA Compliance
    4. Prescriber 360

      Senior Data Scientist

      September 2021 - July 2022

      India

      • Optimized attribution and marketing mix models (MMMs & MTAs) to evaluate channel effectiveness, improving ROI by 15–20% through regression-driven media allocation.
      • Enhanced user journey modeling and conversion tracking, increasing acquisition efficiency by 12% via predictive funnel insights.
      • Developed hybrid LSTM time-series models to forecast campaign performance and engagement.
      • Automated KPI monitoring pipelines for QPS, LTV, and funnel analysis using Spark and Delta Lake.
      • Led enterprise code review initiatives, reducing defects by 80% and standardizing development practices across teams.
      • PySpark
      • Delta Lake
      • LSTM Time-Series
      • Marketing Mix Modeling (MMM)
      • MTA
      • Attribution Models
    5. Cognizant

      Senior Data Scientist

      October 2019 - September 2021

      India

      • Optimized two-tower recommendation architecture for personalized product search, increasing ranking precision and CTR by 18%.
      • Built scalable ML pipelines on AWS (SageMaker, Glue, Airflow, S3), cutting model deployment time by 40%.
      • Enhanced large-scale ranking pipelines on AWS EMR, reducing product search latency by 25% through scoring algorithm optimization and LSH acceleration.
      • Implemented semantic search models using Word2Vec and custom NLP embeddings, improving query relevance and diversity by 15–20%.
      • Developed fairness-aware post-processing layers to balance promoted product visibility while maintaining top-ranking relevance.
      • AWS (SageMaker, Glue, EMR, S3)
      • Airflow
      • Two-Tower Recommender
      • Semantic Search
      • Word2Vec
      • NLP
    6. Tata Consultancy Services (TCS)

      Data Engineer

      December 2015 - September 2019

      India

      • Developed rule-based AML risk scoring modules using Hadoop and HiveQL, improving suspicious activity detection accuracy and reducing false positives by 30%.
      • Engineered end-to-end pipelines for feature extraction from behavioral, temporal, and network transaction data.
      • Migrated legacy SQL systems to cloud to enable scalable Anti-Money Laundering (AML) monitoring and faster SAR investigation workflows.
      • Hadoop
      • HiveQL
      • Big Data Pipelines
      • AML Risk Scoring
      • SQL Migration
    7. ITC Infotech

      Associate Data Warehouse Engineer

      July 2013 - November 2015

      India

      • Optimized large-scale SQL workloads through schema redesign, materialized views, and indexing, improving query performance by 25–40% and saving 1,800+ analyst hours annually.
      • Implemented workload management and caching strategies to prioritize critical jobs and enhance resource utilization by 20%.
      • SQL Server
      • Data Warehousing
      • Query Optimization
      • Materialized Views
      • Performance Tuning

    Research

    Generative AI Application Screening System

    Graduate Research Assistant · Northeastern University

    January 2023 - August 2024

    Boston, MA, USA

    Dean's Commendation

    • Engineered a Generative AI system for graduate application screening, automating evaluation workflows and reducing manual review efforts by 95%, earning commendation from the Dean.
    • Architected a hybrid recommendation framework combining content-based and collaborative filtering approaches to predict applicant–program fit.
    • Integrated Vision Transformers (ViT) and LLMs to analyze multimodal inputs such as transcripts and statements of purpose.
    • Implemented LayoutLMv3 for layout-aware entity extraction from academic transcripts and fine-tuned BERT for accurate subject/grade recognition.
    • Fused textual and visual embeddings using multimodal attention layers, enhancing classification recall by 18%.
    • Deployed production-grade pipelines on GCP (Vertex AI, Cloud Run) achieving high availability and scalable inference across multiple admission cycles.
    • Mentored and led a team of 5 research associates under the Associate Dean’s guidance.
    • GCP (Vertex AI, Cloud Run)
    • LayoutLMv3
    • BERT
    • Vision Transformers (ViT)
    • LLMs
    • Multimodal Attention

    Education

    Master of Science in Data Science

    Northeastern University

    09/2022 – 08/2024 · Boston, MA, USA

    GPA: 4.0 / 4.0 (Dean's Commendation Award)

    Bachelor of Technology in Electronics and Communications Engineering

    West Bengal University of Technology

    08/2009 – 06/2013 · India

    First Class with Distinction

    Certifications

    • AWS Certified Solutions Architect – Associate

      Amazon Web Services

    • AWS Certified Developer – Associate

      Amazon Web Services