Hi, I'm George

CS student at the University of Michigan. I build scalable systems and ship real products and features.

About

I'm an undergraduate at the University of Michigan pursuing a B.S.E. in Computer Science. I've built production systems as a founding engineer, worked in bank tech, shipped an LLM tool across 26 sites, and published first-author research. Outside of code, I'm into philosophy, game design, bouldering, skateboarding, golf, guitar, and calisthenics.

Experience

Building and scaling production systems.

Capital One logo

Capital One

Software Engineer Intern

Jun 2026Aug 2026

8 monthsof data staleness eliminated
  • Delivered an end-to-end production asset certification feature, replacing manual outreach across 2,000+ assets with a secure PUT endpoint under three-tier authorization and matching React components.
  • Built 2 EventBridge-scheduled Lambdas that automate certification reminder emails and daily ServiceNow CMDB sync, eliminating up to 8 months of data staleness from a prior one-time load.
  • Authored idempotent Flyway PostgreSQL migrations that add a CMDB caching layer and tune SQL for compliance reporting.
  • Configured CI/CD across dev, QA, and prod with Jenkins, IAM roles, SecretsManager, and CloudFormation stacks.
AWS Lambda
EventBridge
Python
PostgreSQL
React
Jenkins
BoilerVault Storage LLC logo

BoilerVault Storage LLC

Founding Engineer · Contract

Jan 2026Jun 2026

$130K+revenue reconciled, 0 manual entry
  • Reconciled out-of-order payments through webhook pipelines across 4 Stripe accounts and 3 WordPress sites, processing $130K+ in revenue.
  • Automated unmatched-charge reconciliation, saving 3+ hours of manual work per week by storing and sweeping oldest-first.
  • Built a multi-tenant FastAPI/PostgreSQL/Next.js platform for a 3-campus storage business, with JWT auth across 20+ REST endpoints and a 318-test backend suite.
  • Migrated 475 legacy bookings into 2,000+ records across 5 tables, replacing a failing Zapier Google Sheets workflow.
FastAPI
PostgreSQL
Next.js
TypeScript
Stripe
Nexteer Automotive logo

Nexteer Automotive

Software Engineer Intern

May 2025Aug 2025

87.5%less code-review time
  • Built an LLM-powered IDE extension in Python and TypeScript that parses 300+ internal engineering guidelines and automates compliance checks across C/H files via Azure AI and Claude APIs.
  • Tuned prompt pipelines and few-shot strategies to 95% violation-detection accuracy on embedded steering code.
  • Deployed to 26 sites, saving engineers an estimated 8 hours of manual code review per week.
Python
TypeScript
Azure AI
LLM
Prompt Engineering
Villanova University logo

Villanova University

Data Engineer

Jun 2023Sep 2023

28,000+pathogen isolates analyzed
  • Built R data pipelines that analyze 28,000+ pathogen isolates across 10+ years with PCA and clustering.
  • Published the findings as first author in Antibiotics (2023).
R
PCA
Clustering
Data Pipelines

Projects

A mix of systems, ML, and full-stack work. Here are a few I'm proud of.

Architecture
0.482

PR-AUC vs. 0.421 baseline

MeloChron

A causal self-attention encoder summarizes a listener's track history and feeds a 468K-parameter 2-layer MLP that predicts whether they return to a newly heard track. Trained in PyTorch on a cloud RTX 4090, cutting a multi-day run to hours while scaling training data 7x. It beat a strong 0.421 PR-AUC baseline to reach 0.482, a gain confirmed by a paired bootstrap significance test over 100,000 listeners.

PyTorch
Transformers
Self-Attention
Recommender Systems
Audio Embeddings
Preprocessing
0.97

validation AUROC

Image Classification & Transfer Learning

A CNN and a from-scratch Vision Transformer, both built in PyTorch, classify dog breeds. A two-stage transfer-learning pipeline pretrains on a 10-class breed task, then fine-tunes a frozen convolutional backbone on a binary task, reaching 0.93 accuracy and 0.97 AUROC.

PyTorch
CNN
Vision Transformer
Transfer Learning
Architecture
3,000+

Wikipedia docs indexed

Scalable Search Engine

A multi-stage MapReduce pipeline indexes 3,000+ Wikipedia documents, computing TF-IDF scores and per-document normalization factors across a parallel, multi-job architecture. A Flask API serves 3 partitioned index segments, ranking results by PageRank-weighted cosine similarity, and is deployed to AWS.

Python
MapReduce
Flask
TF-IDF
PageRank
100+

users served

Resume Screener

An NLP pipeline built with TF-IDF and a scikit-learn LinearSVC classifies 12,000+ resumes into 43 categories at 82%/94% top-1/top-3 accuracy. RAG retrieval with FAISS and an LLM ranks resumes against live job postings, deployed via Streamlit to 100+ users.

Python
scikit-learn
LinearSVC
RAG
FAISS
Streamlit

Skills

Languages
PythonC/C++JavaSQLTypeScript/JavaScriptHTML/CSS
ML & Frameworks
PyTorchTensorFlowscikit-learnReactFlaskNode.js
Tools
GitDockerAWSPostgreSQLUnix

Education

University of Michigan

B.S.E. Computer Science

GPA 3.81

2024May 2028

Research

Let's talk

Open to Summer 2027 internships in ML and software engineering. The fastest way to reach me is email.

© George Gu