I turn messy, real-world data into models that make decisions.
At BlackBox I cut document retrieval time 98% across 50,000+
documents; at PwC I hit 92% forecast accuracy on $20M/month in
petrochemical trades. My focus is statistical rigor and
pipelines built to be reproducible, not just to work once.
From raw data to a monitored, production endpoint.
Ingest
Raw data from APIs, warehouses & streams
Clean & Engineer
Handle nulls, build features
Train & Validate
XGBoost, Random Forest, cross-validation
Evaluate
Score against holdout & business metrics
Deploy & Monitor
Serve predictions, track drift
Projects
Filter by discipline, or see everything at once.
EquiSight
Data Science
Challenge: walk-forward validated ARIMA vs. gradient-boosted
(XGBoost) forecasting across 59 equities — 33.6% RMSE improvement
at p ≈ 3×10⁻⁹ — then served it through a FastAPI + PostgreSQL
backend and a React/Plotly dashboard, deployed live with
sub-second API response times.
Challenge: built the eval harness and a hand-verified 42-question
golden set before writing a single retriever, then ran a 3-stage
retrieval ablation (BM25 → dense → RRF hybrid fusion) over 109
SEC filings, NAIC model laws, and state insurance bulletins,
scored with bootstrap confidence intervals. Hybrid fusion led on
every metric (recall@10 0.767, MRR 0.495, nDCG@10 0.556);
refusal accuracy held at 91.7%, and hand-tracing every false
refusal showed all 9 were retrieval-recall gaps, not generation
failures.
Challenge: built an NLP pipeline (TF-IDF, t-SNE, K-means, LDA) to
cluster 32,000+ CORD-19 research papers, preserving 95% variance
while reducing a multi-thousand-paper corpus into navigable
topic groups. (Lumiere Education Research)
Challenge: cut document retrieval time 98% (30 min to <1 min)
across 50,000+ documents with a statistical ranking pipeline,
and automated resume fitment scoring for ~500 applications/cycle
with a TF-IDF-based NLP model — reducing screening time 60%.
PythonNLPTF-IDFStatistical Scoring
Proprietary — built during internship at BlackBox
Petrochemical Price Forecasting
Data Science
Challenge: hit 92% forecast accuracy on $20M/month in
petrochemical (Paraxylene) trades — cutting forecast error from
45% to 28% with XGBoost + Random Forest — then containerized the
model with Docker and deployed on AWS SageMaker, scaling
real-time inference to 10,000+ records/day.
XGBoostscikit-learnDockerAWS SageMaker
Proprietary — built during internship at PwC
Contact
Based in New York, NY — open to Data Science / AI-ML roles.