See what the model changes.


The comparison makes edges, texture, and color differences easy to inspect. A sharper-looking result is not necessarily a more accurate record of the original scene: generative reconstruction can introduce plausible detail.
One model reconstructs. Another challenges it.
- 64 × 64Low-resolution input
- Residual generator16 residual blocks
- Upsample twiceTwo 2× stages
- 256 × 256Reconstructed image
The generator combines convolution, batch normalization, PReLU activation, and skip connections. Sixteen residual blocks work on the image representation before two upsampling stages expand it to four times the input width and height.
The notebook names these stages “pixel shuffler” blocks, but the published implementation uses UpSampling2D after a convolution. That is an implementation difference from sub-pixel rearrangement, and part of what makes this a learning study rather than an exact reproduction.

Optimize for more than matching pixels.
The original SRGAN paper combines adversarial loss with a feature-based content objective. Its motivation is that minimizing pixel error alone can favor smooth, perceptually unconvincing output. Read Ledig et al.’s research.
My implementation uses an ImageNet-pretrained VGG19 feature extractor alongside a convolutional discriminator. The discriminator learns to distinguish generated images from high-resolution examples, while the generator is trained against both realism and feature similarity.
The repository contains a 15-epoch training-loop configuration and checkpoints from different runs. I present the saved visual evidence without treating those files as proof of one continuous, fully documented experiment.
The output is evidence. So are its imperfections.

The project demonstrates a complete adversarial training pipeline and visible differences between input and reconstructed images. It does not publish a reproducible held-out PSNR, SSIM, or perceptual benchmark alongside these examples.
A stronger follow-up would compare bicubic interpolation, a reconstruction-only model, and the adversarial model on identical held-out inputs. I would report quantitative scores together with visual failure cases, especially color shifts and fabricated textures.
The hardest constraint: limited training compute.
The hardest part was training a generative model with limited compute. Super-resolution involves a generator, a discriminator, and a perceptual objective, so every experiment has a cost beyond simply increasing image resolution.
The experience made the compute budget a practical part of the engineering problem. For a follow-up, I would prioritize controlled comparisons and reusable checkpoints so each training run answers a specific question about image quality.