An end-to-end deep learning system that analyzes visual input, extracts high-level semantic features using a Convolutional Neural Network (CNN), and generates natural language descriptive captions using a Recurrent Neural Network (LSTM).
- Automated visual feature extraction using a pre-trained CNN backbone
- Sequence generation pipeline using LSTM networks trained with word tokenization
- Greedy search and beam search decoding strategies for caption generation
- Interactive web-based user interface for uploading custom images and previewing captions
- Comprehensive Jupyter Notebook documenting data preprocessing, vocabulary building, and model training
Input Image
↓
CNN Feature Extractor (Encoder)
↓
Extracted Visual Features (Embedding)
↓
Tokenized Word Sequence (Decoder Context)
↓
LSTM Language Model
↓
Caption Prediction / Output Text
IMAGE-CAPTIONING-DEEP-LEARNING/
│
├── images/
│ ├── Block Diagram.png
│ ├── frontend.jpg
│ ├── Test1.png
│ ├── Test2.jpg
│ ├── DL_Proj_Image_1.jpg
│ ├── DL_Proj_Image_2.jpg
│ ├── DL_Proj_Image_3.jpg
│ └── DL_Proj_Image_4.jpg
│
├── notebook/
│ └── DL_Project.ipynb
│
├── .gitignore
├── README.md
└── requirements.txt
| Input Test Image | Generated Caption Output |
|---|---|
![]() |
Model-generated descriptive caption |
![]() |
Model-generated descriptive caption |
git clone [https://github.com/umerharoon890/image-captioning-deep-learning.git](https://github.com/umerharoon890/image-captioning-deep-learning.git)
cd image-captioning-deep-learningpython -m venv .venvFor Windows PowerShell:
.\.venv\Scripts\Activate.ps1For macOS/Linux:
source .venv/bin/activatepip install -r requirements.txtTo inspect model training, evaluation, and tokenization:
jupyter notebook notebook/DL_Project.ipynbTo launch the web interface:
streamlit run app.pyBridging computer vision and natural language processing is fundamental for accessibility tools, automated visual documentation, and media indexing. This project demonstrates how multimodal deep learning architectures extract representations from convolutional layers and map them directly to natural language syntax.



