Data Engineer @ Melio | MS Data Science, Rutgers University π Sunnyvale, CA
I build end-to-end data pipelines, ETL/ELT workflows, and production ML systems β currently focused on large-scale biological data processing, computer vision for instrumentation, and cloud-native analytics infrastructure.
- Production data pipelines for multi-chip biological assay analysis β melt-curve processing, clustering, classification, and QA automation.
- Computer vision + signal processing for well-detection and chip-fill analysis, supporting multiple new chip generations.
- Modular ETL workflows on AWS with Python + Docker, with a focus on reproducibility, observability, and runtime optimization.
- Anomaly / novelty detection research using One-Class SVM, LOF, and open-set classification on time-series data.
Languages
Python SQL R Java Bash
Data & ML
Pandas NumPy scikit-learn XGBoost LightGBM CatBoost TensorFlow PyTorch OpenCV
Data Engineering
ETL/ELT Parquet Data Modeling Workflow Orchestration PostgreSQL MySQL MongoDB SQLite
Cloud & DevOps
AWS (ECS, S3, Elastic Beanstalk) Azure (Blob Storage, Web Apps) GCP (Vertex AI) Docker GitHub Actions CI/CD
Analytics
Tableau Matplotlib Seaborn Plotly
Ensemble pipeline (CatBoost + LightGBM + XGBoost) for credit-risk classification. Containerized inference service deployed to AWS ECS, Azure Web Apps, and AWS Elastic Beanstalk with full CI/CD via GitHub Actions.
Python XGBoost Docker AWS Azure GitHub Actions
Automated weekly scraping of football statistics with BeautifulSoup, persisted to SQLite, exported to Parquet, and published to Azure Blob Storage on a GitHub Actions schedule.
Python BeautifulSoup SQLite Parquet Azure
Survival analysis with Kaplan-Meier curves on imputed clinical datasets (MICE, PMM) to identify significant predictors of patient outcomes.
R Python Survival Analysis MICE
Hybrid search backed by MySQL (relational) and MongoDB (NoSQL) with query optimization and caching for low-latency retrieval.
MySQL MongoDB Python
Cleaned and transformed Uniform Crime Reporting data; applied regression models and hypothesis testing to surface state-level trends.
R Statistics Regression
- Distributed data processing (Spark, Airflow) and lakehouse architectures
- MLOps best practices β model monitoring, drift detection, feature stores
- LLM evaluation and applied generative AI for analytics workflows
- πΌ LinkedIn
- π GitHub
- βοΈ manassarthak@gmail.com
