Diffusion Systems Experiment
Stable Diffusion v1.5: FP32 vs FP16 vs Gradient Checkpointing
Fine-tuned Stable Diffusion v1.5 on Naruto image-caption data and structured the work as a controlled
systems comparison rather than a one-off model run.
- Shared setup across three runs to isolate precision and memory configuration changes.
- Compared FP32 baseline, FP16 mixed precision, and FP16 with gradient checkpointing.
- Used a fixed inference seed, prompt suite, guidance scale and diffusion-step count for comparison.
- Published each run as a separate Hugging Face Diffusers pipeline with model documentation.
Stable DiffusionDiffusersFP16
Gradient CheckpointingTensorBoard
Custom Generative Model
Few-shot UNet2D DDPM at 256 × 256
Built and trained a custom unconditional diffusion model with Hugging Face Diffusers using a six-stage
UNet2D architecture, attention blocks, BF16 mixed precision and a 1,000-step DDPM scheduler.
- Trained on a small 272-example dataset to explore behavior in a constrained-data regime.
- Used batch size 8 with gradient accumulation 4 for an effective batch size of 32.
- Tracked learning-rate warmup, loss dynamics, throughput, periodic sampling and checkpoint cadence.
- Published the resulting pipeline and a reproducibility-focused Hugging Face model card.
UNet2DModelDDPMBF16
Few-shotPyTorch
Parameter-Efficient LLM Fine-Tuning
BLOOM-7B1 LoRA with PEFT
Fine-tuned BLOOM-7B1 using LoRA adapters on an English quotes dataset, freezing the base model and
updating only a small adapter parameter set.
- 7,864,320 trainable parameters out of 7,076,880,384 total, roughly 0.11%.
- Used FP16, gradient checkpointing, batch size 4 and gradient accumulation 4.
- Completed 200 optimizer steps with a reported aggregate training loss of 2.3122.
- Stored the resulting adapter as an approximately 31.5 MB Safetensors artifact.
TransformersPEFTLoRA
BLOOM-7B1FP16
Foundational Diffusion Work
UNet Diffusion Baselines and Architecture Iteration
Progressed from a small diffusion baseline to modified UNet experiments before moving into the larger
few-shot and Stable Diffusion studies. These notebooks focus on understanding the denoising process,
scheduler behavior, architecture changes and generated samples from first principles.
UNetNoise SchedulingDenoisingPyTorch