19

Automatic Image Captioning System

Automatic Image Captioning System

Automatic Image Captioning System brings together the realms of computer vision and natural language processing to allow machines not just to see images, but to describe them in natural language. At the heart of the system lies an encoder-decoder architecture: a Convolutional Neural Network (CNN) pre-trained on large-scale datasets serves as the encoder, extracting nuanced visual features, while a Long Short-Term Memory (LSTM) network acts as the decoder, generating coherent captions word by word. Inspired by the seminal “Show and Tell: A Neural Image Caption Generator” paper, this implementation follows a proven yet adaptable model, designed to take an input image and produce a descriptive sentence that captures its content. The repository is structured thoughtfully: you’ll find Jupyter notebooks for training and experimentation, a well-organized Conda environment setup (environment.yml), scripts for model training (train_utils.py), data handling (data.py, utils.py), and a tracking framework using Weights & Biases and Ax for hyperparameter tuning.

The captioning pipeline begins with preparing the Flickr8k dataset, one of the classic image-caption datasets, organizing image files and paired captions into training, validation, and test splits. During training, the CNN encoder transforms each image into a compact and informative feature vector, which the LSTM receives at the start of a sequence generation process. The decoder then iteratively builds a sentence, conditioned on both these features and the previously generated words. A suite of hyperparameters (learning rate, embedding dimensions, LSTM layers) is tuned via Ax and the training process tracked using Weights & Biases, providing transparency and reproducibility. According to the README, while the original paper reports a BLEU score of 27.2, this model achieves a modest, but valuable, BLEU score of around 11 on its test split, indicating the system generates captions with reasonable fidelity, especially considering domain differences in training data. Model speed is also noted: after initial loading, caption generation takes around 5 seconds per image, making the system practical for demo or prototype use.

The value of this project lies in its technical implementation and also in its accessibility, modularity, and potential for expansion. Packaged under an MIT license, it’s primed for experimentation. Researchers can swap in upgraded encoders (like Transformer-based vision models), tune LSTM parameters, or expand to richer datasets like MS COCO to improve performance. Its documentation and notebook-based approach make it educational as well as practical, perfect for users exploring image–language models from end to end. Potential extensions include adding attention mechanisms (Show, Attend, and Tell), beam search for more fluent captions, or even transformer decoder architectures for better sequence modeling. In real-world terms, this pipeline could power accessibility tools that describe images for visually impaired users, automate content tagging for image databases, or empower developers with plug-and-play visual understanding modules. This project stands out as a robust foundation in image captioning, both a learning tool and a launch pad for richer, more capable systems.

Faculty

Students

  • Muhammad Abdullah
  • Usama Athar
18

Real-Time Traffic Density Estimation

Real-Time Traffic Density Estimation

In this project, we use YOLOv8 to transform raw video feeds into actionable insights about traffic volumes. Traditional traffic monitoring systems often rely on static sensors (loop detectors or radar) that can miss dynamic patterns or require extensive infrastructure. This project takes advantage of affordable camera setups and deep learning to detect vehicles with precision and count them live in defined areas of interest. By fine-tuning a YOLOv8 model on a custom vehicle dataset, the system not only identifies vehicles but translates their movements into density maps. These maps can show congestion hotspots and temporal flow changes, providing real-time situational awareness to traffic engineers, urban planners, or smart city systems.

This work goes through loading a pre-trained YOLOv8n model, fine-tuning it with custom images and annotations, and then deploying it on a sample video. YOLOv8’s architecture allows for fast, high-accuracy detections even on edge devices, thanks to its efficient backbone and inference pipeline. Once vehicles are detected in each frame, their positions are tallied within predefined regions (such as lanes or junction zones) and translated into vehicle density statistics. The accompanying sample video and resulting traffic density analysis images demonstrate how the system highlights dense traffic clusters and tracks peaks over time. Performance metrics like training loss curves and confusion matrices ensure users can assess model accuracy and fine-tune further as needed. This modular architecture means you can swap in different datasets or expand the system to track vehicle types or multi-camera feeds.

The real magic of this system lies in its practicality and adaptability. It is a demonstration of how vision-based AI can make traffic insights more accessible, scalable, and responsive. By offering a full pipeline, from dataset preparation (image folders, YAML configs) to training scripts, model outputs, and inference demos, this work empowers analysts, developers, and city teams to deploy end-to-end traffic monitoring with minimal setup. From congested intersections to dynamic toll plazas or even parking lots, this solution scales from single-lane monitoring to complex multi-camera networks. Licensed under MIT and well-documented, it invites collaboration and extension, perhaps integrating historical density trends, alerting systems, or traffic signal optimization modules. In an age of urban mobility challenges, this project stands out as a ready-to-use blueprint for smarter, data-driven traffic management.

Faculty

Students

  • Labib Kamran
  • Usama Athar
Gemini_Generated_Image_jxwstyjxwstyjxws

AgroData: Largescale Early / In-Season Crop-Type Mapping

AgroData: Largescale Early / In-Season Crop-Type Mapping

Pakistan’s agricultural sector, vital to its economy and the livelihood of millions, faces critical inefficiencies exacerbated by climate change and poor management. Despite being a leading producer of wheat, rice, and cotton, the country experiences significant crop wastage, up to 15% for wheat and over 20% for fruits and vegetables, due to inadequate planning, outdated data collection methods, and ineffective supply chain management. This inefficiency compromises local food availability and disrupts the balance between food imports and exports, leading to economic instability and increased reliance on costly imports. The primary issue is the absence of a centralized, up-to-date agricultural data system. The lack of real-time data hampers the optimization of crop production, resource management, and strategic decision-making, resulting in resource wastage and threats to national food security.

ad1

This project introduces an advanced large-scale framework for early and in season crop-type mapping to enable crop yield prediction in Pakistan. There are significant threats to food security posed by population growth and climate change, which demand timely and reliable crop data for yield prediction and agricultural planning. Traditional methods relying on field surveys are slow and subjective. This project leverages quasi real-time satellite imagery and advanced deep learning methods to provide valuable crop-type data. The framework includes a distributed data hub, user-friendly access, AI-powered analytics, and predictive insights to support decision-making. Key components feature the AI prediction model, data ingestion pipeline, a frontend interface, and a predictive analytics dashboard. The aim is to visualize maps of different regions of Pakistan labeled by crop type and to enable yield estimation, monitoring of agricultural patterns, and assessing weather impacts, addressing food security and broader environmental concerns.

ad2
ad3

Faculty

Students

  • Muhammad Umer Khan
  • Shalina Riaz
  • Syed Hashir Ahmad Kazmi
FinTech Frontier (2)

05 Day Master Class Workshop On Leveraging Analytics, AI and Data Sciences

ICESCO Chair of AI and Data Analytics for Business in Collaboration with Pakistan-U.S. Alumni Network (PUAN) at NUST SEECS.

A 05 Day Workshop on

Leveraging Analytics, AI and Data Sciences

July 8th-12th, 2024NUST

About The Workshop

Funded by the U.S. Mission to Pakistan, the Pakistan-U.S. Alumni Network (PUAN) in collaboration with ICESCO Chair of AI and Data Analytics for Business and in partnership with NUST, Islamabad, successfully concluded a five-day Masterclass on ‘Leveraging Analytics, AI, and Data Sciences‘ from July 8-12. This masterclass featured a diverse group of dozens of alumni from various backgrounds and offered comprehensive training with interactive lectures by sectoral experts. The learning approach included case studies centered around various sub-themes, group discussions, and presentations.

 

Workshop 2024

Agenda

1

Day 1: Exploration of secure cyber policies and practices to promote entrepreneurship, featuring a case study from Food Panda.

2

Day 2: Delving into cybersecurity and data protection with practical applications.

3

Day 3: Demystification of Artificial Intelligence and its impact on tech businesses.

4

Day 4: Learning about data science and analytics to gain valuable business insights.

5

Day 5: Attending a captivating panel discussion on harnessing AI, data analytics, and data science for innovation and transformation.

 

Expected Outcomes

Enhanced understanding of secure cyber policies and their role in entrepreneurship.

Practical knowledge and applications in cybersecurity and data protection.

A clear grasp of Artificial Intelligence concepts and their business impacts.

Skills in data science and analytics to derive valuable business insights.

Insights into the use of AI, data analytics, and data science for innovation and transformation.

 

Target Audiences

1

Professionals in technology and data-related fields

2

Entrepreneurs looking to leverage data sciences and AI

3

Cybersecurity experts

Resource Person

Prof Dr. Muhammad Moazam Fraz

Director ICESCO Chair
Professor and HoD ( AI & Data Science Dept.)
SEECS, NUST

Participants

0 attendees

Testimonials

“"This masterclass provided a perfect blend of theoretical knowledge and practical applications. The Food Panda case study was particularly enlightening."”
Samira KhalidBusiness Analyst, TechDynamics
“"The sessions on cybersecurity and data protection were very hands-on and immediately applicable to our work."”
Ahmed ZafarData Scientist, InnovateCorp
“"AI is no longer a mystery to us! The expert explanations and business implications discussed were extremely valuable."”
Maria KhanStrategic Planner, BizGrowth Solutions
“"Data science and analytics have become much clearer to me. I can now see how to apply these insights to drive business decisions." ”
Omar FarooqHead of Analytics, FutureEnterprises

Event Pictures

Digital Presence and Online Visibility

1

Shared on NUST LinkedIn

Click Here
2

Shared on SEECS Event Page

Click Here

NUST: Facebook

SEECS: Facebook

MachVIS: Facebook

ICESCO Chair: Twitter

ICESCO Chair: Facebook

Gemini_Generated_Image_tnh9botnh9botnh9

AI-Powered Cephalometric Landmark Detection for Orthodontic Diagnosis: Framework, Dataset, and Detection Challenge

Quantitative cephalometric analysis is a standard clinical and research tool in modern orthodontics which plays an integral role in orthodontic diagnosis, maxillofacial surgery, and treatment planning. The accurate identification and reproducible localization of cephalometric landmarks allows the quantification and classification of anatomical abnormalities. The traditional manual way of marking cephalometric landmarks on lateral cephalograms is a very time-consuming job and is miles hard to achieve stable detection accuracy because of uneven professionalism of orthodontists. Endeavors to develop automated landmark detection systems have persistently been made but they are inadequate for clinical orthodontic applications because of low reliability of specific landmarks.

Gemini_Generated_Image_34tlav34tlav34tl

An Attention-Driven Hybrid Network for Survival Analysis of Tumorigenesis Patients Using Whole Slide Images

An Attention-Driven Hybrid Network for Survival Analysis of Tumorigenesis Patients Using Whole Slide Images

Survival analysis of cancer patients using Whole Slide Images (WSIs) is crucial in the field of medical statistics, as it helps identify key prognostic factors related to mortality and disease recurrence. However, extracting survival-relevant features from these images is a resource-intensive task and poses computational challenges. Instead of using the complete WSIs, most of the existing models rely on labor-intensive and time-consuming annotations by pathologists to extract features. Alternatively, many models take a selective approach by focusing on key patches from WSIs that potentially can miss important morphological details. Furthermore, most of these approaches employ Convolutional Neural Networks (CNNs) pre-trained on ImageNet to extract survival-related information. However, it’s important to note that there are significant differences in data distribution and domain-specific characteristics when comparing natural images to medical images. Our proposed methodology addresses these challenges by leveraging the advantages of contrastive learning and a transformer-based model. This architectural choice allows us to efficiently process and learn from the data, enabling us to capture intricate spatial relationships and contextual information within image patches. This capability is crucial for understanding the complex and heterogeneous nature of pathological tissues. Additionally, to make the most of the available information, we employ a strategy where we randomly select and utilize features extracted from 8,000 patches for each patient. This random selection process is carefully designed to encompass a diverse and relevant set of information from the patient’s dataset, ensuring a comprehensive representation of the input data within our transformer model. We evaluated the performance of our model using the TCGA-GBM and TCGA-LUSC datasets, and it achieved a concordance index (c-index) of 0.7892 and 0.744, respectively.

Video

Faculty

Students

  • Arshi Parvaiz

Publications

  • S. Nasir, A. Parvaiz, M. M. Fraz. “Nuclei and glands instance segmentation in histology images: a narrative review”, In Artificial Intelligence Review 56 (8), 7909-7964 (2023) https://doi.org/10.1007/s10462-022-10372-5
  • Parvaiz, M. A. Khalid, R. Zafar, H. Ameer, M. Ali, M. M. Fraz. “Vision Transformers in medical computer vision—A contemplative retrospection”, In Engineering Applications of Artificial Intelligence 122, 106126 (2023) https://doi.org/10.1016/j.engappai.2023.106126
  • A. Parvaiz, E. S. Nasir, M. M. Fraz. “From Pixels to Prognosis: A Survey on AI-Driven Cancer Patient Survival Prediction Using Digital Histology Images”, In Journal of Imaging Informatics in Medicine, 1-24 (2024) https://doi.org/10.1007/s10278-024-01049-2