Machine Learning Engineering Portfolio

FROM MODEL TRAINING
TO SYSTEMS TRADEOFFS

A practical portfolio of deep learning experiments built in PyTorch and the Hugging Face ecosystem. The work moves from tensor fundamentals into diffusion model training, Stable Diffusion fine-tuning, precision and memory optimization, and parameter-efficient adaptation of a 7B language model. The focus is not only on getting models to run, but on understanding architecture choices, reproducibility, inference behavior, checkpointing, and deployment-ready artifacts.

7+technical notebooks spanning fundamentals, diffusion, Stable Diffusion and PEFT
6published model artifacts across diffusion and language modeling experiments
7Bparameter base model adapted with LoRA while training ~0.11% of parameters
3-waycontrolled Stable Diffusion comparison across FP32, FP16 and gradient checkpointing

Case Studies

Diffusion Systems Experiment

Stable Diffusion v1.5: FP32 vs FP16 vs Gradient Checkpointing

Fine-tuned Stable Diffusion v1.5 on Naruto image-caption data and structured the work as a controlled systems comparison rather than a one-off model run.

  • Shared setup across three runs to isolate precision and memory configuration changes.
  • Compared FP32 baseline, FP16 mixed precision, and FP16 with gradient checkpointing.
  • Used a fixed inference seed, prompt suite, guidance scale and diffusion-step count for comparison.
  • Published each run as a separate Hugging Face Diffusers pipeline with model documentation.
Stable DiffusionDiffusersFP16 Gradient CheckpointingTensorBoard
Custom Generative Model

Few-shot UNet2D DDPM at 256 × 256

Built and trained a custom unconditional diffusion model with Hugging Face Diffusers using a six-stage UNet2D architecture, attention blocks, BF16 mixed precision and a 1,000-step DDPM scheduler.

  • Trained on a small 272-example dataset to explore behavior in a constrained-data regime.
  • Used batch size 8 with gradient accumulation 4 for an effective batch size of 32.
  • Tracked learning-rate warmup, loss dynamics, throughput, periodic sampling and checkpoint cadence.
  • Published the resulting pipeline and a reproducibility-focused Hugging Face model card.
UNet2DModelDDPMBF16 Few-shotPyTorch
Parameter-Efficient LLM Fine-Tuning

BLOOM-7B1 LoRA with PEFT

Fine-tuned BLOOM-7B1 using LoRA adapters on an English quotes dataset, freezing the base model and updating only a small adapter parameter set.

  • 7,864,320 trainable parameters out of 7,076,880,384 total, roughly 0.11%.
  • Used FP16, gradient checkpointing, batch size 4 and gradient accumulation 4.
  • Completed 200 optimizer steps with a reported aggregate training loss of 2.3122.
  • Stored the resulting adapter as an approximately 31.5 MB Safetensors artifact.
TransformersPEFTLoRA BLOOM-7B1FP16
Foundational Diffusion Work

UNet Diffusion Baselines and Architecture Iteration

Progressed from a small diffusion baseline to modified UNet experiments before moving into the larger few-shot and Stable Diffusion studies. These notebooks focus on understanding the denoising process, scheduler behavior, architecture changes and generated samples from first principles.

UNetNoise SchedulingDenoisingPyTorch

Stable Diffusion Experiment Matrix

RunPrecisionGradient CheckpointingPurpose
sd-naruto-fp32-baselineFP32NoReference configuration
sd-naruto-fp16FP16NoMixed-precision comparison
sd-naruto-fp16-grad-checkpointingFP16YesMemory-oriented configuration

Notebook Progression

Engineering Themes

Model architecture experimentation, controlled training comparisons, mixed precision, gradient accumulation, gradient checkpointing, diffusion schedulers, checkpoint management, PEFT/LoRA, reproducibility, inference testing, model cards, and publishing deployable artifacts to the Hugging Face Hub.