Rankush Vishwakarma
status: active · role: data-engineering · region: noida-in

Rankush Vishwakarma

Senior Data Engineer · Data Team Lead @ Appinventiv

5+ years building enterprise ETL/ELT pipelines and cloud data platforms. Currently leading data engineering across three concurrent enterprise projects — a Microsoft Fabric lakehouse for Coca-Cola's dispenser fleet and a churn-prediction platform processing 900+ tenant accounts among them. Previously at DXC Technology, where a self-initiated reconciliation pipeline validated up to 50,000 insurance policy records a day against a legacy DB2 mainframe migration.

$850/moAzure cost saved
15→2hrsnightly batch runtime
300+discrepancies caught
500+tables migrated

contact — rankush.vishwa@gmail.com · +91 825 301 5353 · LinkedIn · GitHub

experience

// professional experience

  1. Senior Data Engineer · Data Team Lead — Appinventiv

    Aug 2025 → present

    Noida, Uttar Pradesh, India · Hybrid

    • Leading data engineering across 3 concurrent enterprise projects; managing 2 data engineers, 4 data analysts, 1 AI/ML engineer
    • Architecting a Microsoft Fabric medallion pipeline (Bronze/Silver/Gold) for Coca-Cola dispenser telemetry — live on 150 of a planned 1,000+ dispensers, T-1 daily batch cycle
    • Redesigned a client pipeline with Celery task queuing + chunked data movement, cutting daily bug volume from 15–20 to 4–5 over 6 sprints
    • Restructured Azure pod/replica configuration across non-prod environments, cutting cloud spend by $850/month
    • Built a multi-tenant churn-prediction system (RetainIQ) across 900+ tenant accounts — flagged ~$27K in exposure in one tracked month, win-back campaigns recovered ~$5,680
    • Migrated 500+ tables MySQL → PostgreSQL, 30% performance improvement
    • Replaced manual client onboarding with an automated onboarding + migration pipeline, with alerting

    python · sql · airflow · dbt · snowflake · azure · aws · kafka · databricks

  2. Data Engineer & Python Developer — DXC Technology

    Jan 2021 → Aug 2025

    Indore, India · Remote

    • Self-initiated an automated daily reconciliation pipeline after identifying manual validation (20–30 policies/day) couldn't scale to 15,000–50,000 policies/day during a legacy DB2-to-cloud migration — caught 300+ discrepancies manual review missed; still in active use (v15) post-handoff
    • Architected enterprise ETL pipelines across 1,000+ tables, 75% speed improvement, 99.9% reliability
    • Rewrote/re-indexed 250+ SQL queries and restructured a nightly batch job, cutting runtime from 8–9 hours to under 2
    • Built a policy data comparison tool, cutting analysis time from 6 hours to 15 minutes — 180+ hours/month saved
    • Delivered a 15-day Python/ML training program to 500+ engineers, cutting external vendor dependency by 60%
    • Led 2 critical projects with 100% on-time delivery while mentoring 3 junior developers
    • 3× Facilitator Award recipient (2021–2024) for technical leadership and delivery

    python · sql · ssis · ibm db2 · azure data factory · azure sql

// entrepreneurial journey

  1. AceInterview.in — AI Interview Platform

    2022 → 2024

    Founder & Developer · side project

    • Full-stack AI interview platform, 500+ users, 95% satisfaction rate
    • Multi-modal AI assistant with real-time scoring and personalized feedback
    • Resume analyzer with job matching, improving user success rates by 35%

    python · fastapi · gpt-4 · javascript · tailwind · azure · github actions

  2. StressAway — Mental Health Platform

    2022 → 2024

    Co-Founder & Tech Lead · side project

    • 100+ monthly therapy sessions delivered, 87% user satisfaction
    • Conversational AI with emotion-sensitive responses and stress analysis

    python · fastapi · mongodb · docker · openrouter · nlp

  • PulseLog — High-Performance Python Logging Infrastructure

    2026 → present

    Creator · independent engineering project

    • Built a high-performance Python logging system designed for concurrent, high-throughput production workloads
    • Designed asynchronous, buffered logging architecture with emphasis on producer-side overhead and concurrent correctness
    • Benchmarked logging throughput, CPU cost, batching, and concurrency behavior against established Python logging systems
    • Explored performance trade-offs across queueing, serialization, I/O, batching, and multi-threaded workloads

    python · concurrency · async · queues · high-throughput · benchmarking · systems engineering

  • // education

    1. B.Tech, Computer Science — Amity University, MP

      2017 → 2021

      Capstone in applied ML/NLP systems.

    portfolio

    Rendered as a job log, not a grid — status reflects whether the work is live, shipped, or archived, and only rows with a real public artifact are linked.

    active Coca-Cola Fabric Lakehouse data-engineering private active RetainIQ — Churn Prediction & Win-Back Platform data-engineering / ml private active Pulselog data-engineering Public PyPi library shipped AceInterview.in — AI Interview Platform ai-platforms aceinterview.in →
    shipped StressAway — Mental Health Platform ai-platforms no public link
    archived Amazon Alexa Data Analysis nlp github → archived Emotion Extraction nlp github → archived Yelp Rating Prediction nlp github → archived Soccer Player Rating Prediction machine-learning github → archived Wage Class Prediction machine-learning github → archived Boston House Prediction machine-learning github → archived Boston House Prediction — Random Forest machine-learning github → archived Titanic Survival Prediction machine-learning github → archived Women Affair Classification machine-learning github → archived Software Requirement Classification machine-learning github → archived Churn Modelling — ANN deep-learning github → archived Jarvis 1.1 (Extended) ml-ops github → archived Jarvis 1.0 (Base) ml-ops github → archived PM2.5 Air Concentration Prediction time-series github → archived Traffic Time Series Prediction time-series github → archived OML Streaming GRC Time Series Prediction time-series github → archived Demand Per Hour Prediction time-series github → archived Stock Market Prediction time-series github → archived Historical Stock Market Time Series Clustering time-series github →

    certifications

    contact

    Reach out directly — no form, no middleman.