Gemini_Generated_Image_b0z1nvb0z1nvb0z1

HunarmandAI: Culturally-Aware Intelligent Tutoring for Low-Literacy Urdu Speakers

HunarmandAI: Culturally-Aware Intelligent Tutoring for Low-Literacy Urdu Speakers

HunarmandAI is an intelligent tutoring system designed specifically for Pakistani auto electricians who often struggle with limited access to training resources due to low literacy levels and language barriers. While most digital platforms and AI tools are dominated by English or highly formal Urdu, the everyday language spoken in workshops is far more informal, localized, and mixed with English technical terms. HunarmandAI bridges this gap by introducing a bilingual Urdu-English small language model (SLM) that can understand and respond in a way that feels natural to electricians. By embedding specialized automotive electrical knowledge directly into the model, the project creates a culturally and linguistically accessible tutor that works entirely offline, ensuring usability even in areas with little or no internet connectivity.

This project pushes the boundaries of what small language models can achieve by focusing on reasoning, accessibility, and edge deployment. Unlike general-purpose chatbots, HunarmandAI is fine-tuned with instruction-following datasets, step-by-step reasoning examples, and domain-specific content derived from real-world resources used by technicians. Through methods like instruction tuning, chain-of-thought optimization, and reinforcement learning with human feedback, the model is trained not only to provide correct answers but also to explain its reasoning in clear, easy-to-follow steps. This makes the tutor much more than just a Q&A tool. It becomes a practical guide that teaches problem-solving, repair techniques, and diagnostic skills in a way that mirrors how electricians actually work.

HunarmandAI has a strong social and economic impact as well. It empowers auto electricians from lower and lower-middle-class backgrounds to learn independently, improve their craft, and enhance their earning potential without relying on costly training centers or constant internet access. At the same time, it contributes to AI research by demonstrating how small language models can be adapted for reasoning-rich, domain-specific tasks on low-power devices. By releasing the model, dataset, and mobile app openly, the project also lays a foundation for future culturally-aware AI tutors in other vocational fields, advancing digital inclusion and supporting sustainable skill development in underserved communities.

Faculty

Students

  • Muhammad Nabeel
  • Hafiza Adeela Arif
  • Abubakar Khan
Gemini_Generated_Image_l9fsgbl9fsgbl9fs

MedLingo: Medical LLM with Mixture of Experts

MedLingo: Medical LLM with Mixture of Experts

MedLingo is a clinical decision-support system designed to bring specialization and precision into medical AI. Unlike general-purpose language models that try to cover everything but often lack depth in critical areas, MedLingo is built around a Mixture-of-Experts framework. Instead of relying on a single, monolithic model, it coordinates multiple smaller, domain-specific language models—each trained in a specialty such as ophthalmology, neurology, or oncology. This modular setup allows MedLingo to act more like a team of medical experts: when a clinician presents a case, the system routes the input to the most relevant specialists, who then collaborate and reconcile their findings. By integrating multimodal data (clinical text, lab results, medical imaging), the platform delivers context-aware analysis that mirrors the way real doctors synthesize evidence.

This project combines natural language understanding with medical image encoding and structured data processing, ensuring no critical detail is left out of the diagnostic process. The system is evaluated on a wide range of benchmarks, from MIMIC-III and PubMedQA for clinical text to CheXpert and ISIC for imaging, ensuring that its performance is robust across diverse data types. Its architecture is carefully designed for efficiency: the Mixture-of-Experts controller dynamically selects the most relevant models for each query, reducing unnecessary computation and maintaining low latency. This makes MedLingo not only accurate but also scalable, allowing new expert models to be added over time without overhauling the system.

The impact of MedLingo extends beyond technical novelty. For healthcare professionals, it offers faster, more reliable diagnostic support that could reduce errors, accelerate clinical decision-making, and improve patient outcomes. For researchers, it provides a modular blueprint for how AI can be structured in medicine, demonstrating how collaboration between expert models can enhance reasoning in ways a single system cannot. And for the broader community, the project contributes open-access code, datasets, and methodologies, fostering transparency and future innovation in medical AI. By making clinical decision-support systems more specialized, adaptable, and multimodal, MedLingo represents a step toward AI that works not as a black box, but as an intelligent and trusted partner in healthcare.

Faculty

Students

  • Muhammad Samiullah
  • Sarmad Sultan
  • Ahmad Jan
Gemini_Generated_Image_w8kd0gw8kd0gw8kd

LiDAR-Based Plant Phenotyping for Precision Agriculture

LiDAR-Based Plant Phenotyping for Precision Agriculture

This project tackles one of the biggest bottlenecks in modern crop science: how to measure plant traits quickly, accurately, and at scale. Traditional phenotyping is still largely manual. Researchers take rulers and calipers into fields, a process that is slow, destructive, and vulnerable to human error. This project proposes a smarter alternative by harnessing LiDAR sensors to capture rich 3D representations of plants and applying deep learning models to directly interpret that structural data. Instead of reducing plants to simplistic measurements like height or canopy width, the system learns to estimate key biological traits such as above-ground biomass, canopy volume, and other indicators of crop performance. By developing a full software pipeline and integrating both open datasets and custom LiDAR scans, this project aims to create a scalable, non-destructive tool that makes high-throughput phenotyping practical for breeders and agronomists.

The technical backbone of the project is a novel deep learning model, nicknamed AgriFormer, designed to predict above-ground biomass directly from 3D point cloud data. The architecture processes raw LiDAR scans through a multiscale encoder-decoder and then reconstructs dense plant structures to deliver highly accurate predictions. Crucially, this avoids the need for hand-crafted features or voxelization, which limit flexibility and accuracy. Early experiments using open-source datasets show promising results with lower error rates compared to conventional approaches. Alongside this, we build an in-house dataset using custom 2D LiDAR hardware mounted on a UAV platform, generating full 3D reconstructions of target crops.

The significance of this project lies in its potential to reshape how agricultural research is conducted. By lowering the cost and effort of high-fidelity phenotyping, it enables crop scientists to study larger populations with greater accuracy, accelerating breeding programs aimed at higher yields and climate resilience. The methodology also lowers the barrier to entry for labs with limited resources, since the conversion of 2D LiDAR scans into 3D point clouds offers a cost-effective alternative to expensive commercial systems. The pipeline and datasets produced in this work can power a new wave of data-driven agriculture, helping farmers and researchers make smarter decisions about crop management and varietal selection.

Faculty

Students

  • Abdul Wahab
  • Faareh Ahmed
  • Malik Shahzaib Khan

Publications

  • Khan, M.S., Ahmed, F., Wahab, A., Zafar, Z., Berns, K. and Fraz, M.M., 2025, October. AgriFormer: Advancing 3D LiDAR-based Biomass Prediction through Hierarchical Feature Learning. In 2025 IEEE/ACS 22nd International Conference on Computer Systems and Applications (AICCSA) (pp. 1-6). IEEE. https://doi.org/10.1109/AICCSA66935.2025.11315189
Gemini_Generated_Image_j0k42aj0k42aj0k4

GlacioVision: Modeling Glacier Retreat and Water Security in Pakistan Through Remote Sensing Technologies

GlacioVision: Modeling Glacier Retreat and Water Security in Pakistan Through Remote Sensing Technologies

GlacioVision is an intelligent framework to forecast glacier retreat and evaluate its impact on water resources. Pakistan’s glaciers, particularly in the Karakoram and Hindu Kush, serve as vast natural reservoirs that feed the Indus River system and sustain agriculture, hydropower, and drinking water supplies for millions. But rising temperatures and erratic weather patterns are rapidly destabilizing these fragile ice bodies, leading to increased flood risks, seasonal water variability, and long-term freshwater shortages. GlacioVision tackles this challenge by combining multi-year satellite imagery with climate and meteorological data to model how glaciers evolve over time, offering both scientific insights and practical tools for water and disaster management.

At the heart of the project is a deep learning model designed to jointly predict glacier extent and surface elevation. Unlike traditional methods that treat these aspects separately, GlacioVision integrates Sentinel-1 radar imagery, Sentinel-2 optical data, ICESat-2 elevation points, and reanalysis climate variables into a single predictive system. The model leverages convolutional neural networks to extract spatial patterns, long short-term memory modules to capture temporal dynamics, and multilayer perceptrons to embed atmospheric influences. This multimodal, temporally-aware design allows the framework to generate high-resolution predictions of glacier masks and digital elevation models (DEMs), producing a comprehensive view of both surface retreat and volumetric changes. By learning from historical sequences of glacier behavior, the system can forecast near-future changes, providing a critical early warning mechanism for water security planning and climate adaptation.

GlacioVision’s predictive outputs can support government agencies in designing resilient water infrastructure, help disaster managers anticipate glacial lake outburst floods in vulnerable valleys, and guide farmers and hydropower operators in adapting to shifting river flows. The framework’s reliance on open satellite and climate datasets makes it scalable and replicable in other glacier-dependent regions of the world. In Pakistan, however, the project holds especially urgent significance because it provides an accessible way to understand glacier decline.

Faculty

Students

  • Muhammad Sarmad Saleem
  • Abubakar Imran
  • Ali Haider
22

Sports Footage Analysis

Sports Footage Analysis

All major sport teams have always wanted a sports coach’s keen eye and near-perfect memory. This can be embedded in a software system that watches every play, tracks every movement, and extracts meaningful patterns, all in real time. This is what this project sets out to accomplish. Designed for football, the system turns raw game footage into actionable insights for coaches, analysts, and players. It detects and tracks players, the ball, and relevant field markers like lines and goalposts, then computes metrics like player distance covered, speed, possession, heatmaps, and strategic positioning. Far beyond simply recording who’s on the field, this tool transforms video into the kind of deep, data-driven intelligence that can influence game plans, training regimens, and even talent scouting.

The technical architecture underpinning this capability is both modular and powerful. First, the system preprocesses raw footage, aligning frames, calibrating views, and preparing the video for detection. It then employs deep learning object detection to identify players, the ball, and lines. These detections are passed into a tracking module that links objects across frames, creating trajectories over time. From these paths, it computes key metrics, like a player’s total distance run or a ball’s possession time, while overlaying visualizations of movement trails and heatmaps. The results are exportable as CSVs, allowing analysts to combine quantitative data with video playback, making it seamless to correlate critical moments, like a sudden sprint, with tactical decisions. Whether you are measuring how often a striker drops back into midfield or quantifying how possession shifts during a match, the system provides high-resolution visibility into game dynamics.

What makes this project particularly valuable is the bridge it builds between cutting-edge AI and real-world sports analytics. It is not a black-box or prototype, it is a practical, deployable framework. Analysts can use it to automatically generate reports on player performance; coaches can use it to refine strategies; broadcasters can layer live insights during replays. The accuracy of object detection combined with smooth tracking means that patterns emerge clearly, whether highlighting who tended to control the wings, which players made the most high-speed runs, or where formations broke down. By sharing clean, well-structured data alongside visual outputs, the system invites further innovation: one might integrate it with GPS tracking, link it to physiological sensors, or feed its findings into performance dashboards. This project not only brings AI to bear on the complex, fast-moving world of sports, it makes that power usable, transparent, and ready to shape the next generation of athletic performance.

Faculty

Students

  • Abdullah Usama
  • Usama Athar
23

Change Detection in Remote Sensing Imagery

Change Detection in Remote Sensing Imagery

This project targets to detect, localize, and interpret changes in the Earth’s surface through an intelligent transformer-based system. It tackles a vital challenge: automatically identifying regions of change between two images taken at different times (dual-phase imagery), such as areas affected by urban expansion, deforestation, or natural disasters. Traditionally, such detection relies on convolutional methods that struggle with long-range context and subtle temporal differences. By introducing transformer-based architectures, the project brings a fresh perspective, enabling models to compare images not just pixel by pixel, but by capturing spatial and temporal “conversations” between regions across time. This method holds the promise of faster, more accurate mapping of change, key for environmental monitoring, urban planning, and disaster response.

The pipeline for this project is built in PyTorch, implementing transformer models for pixel-level change detection. We adapt and apply models like the Bitemporal Image Transformer (BIT), which first compresses dual-temporal images into meaningful tokens, then uses transformer encoders to model spatial-temporal context and decoders to project these insights back into refined pixel predictions. This approach significantly reduces the compute overhead compared to convolution-heavy systems, while maintaining or surpassing accuracy. The code supports widely-used remote sensing change detection datasets (LEVIR-CD, WHU-CD, DSIFN-CD), providing scripts for training, evaluating, and running demos. Users can recreate the entire workflow, from preprocessing paired images to generating visual change maps. Notably, BIT-based models often outperform standard convolutional baselines using just a third of the parameters and compute cost.

The modular structure of this pipeline means developers can easily experiment with new transformer variants, attention mechanisms, or dataset formats. The integration of dual temporal data opens doors to domain-specific customization. This can be adapted to detect building damage after storms, monitor vegetation loss in wildfire zones, or track glacier retreat. The use of publicly available datasets and open MIT licensing ensures the project can be extended and adapted freely. For practitioners, the ability to visualize and act on accurate change maps, without needing massive compute or complex pipelines, makes this system incredibly practical. This work shows how current AI research can move beyond lab benchmarks into environments where timely and reliable change detection can make a real difference.

Faculty

Students

  • Navaal Iqbal
  • Ayesha Siddiqa
  • Amna Ahmed
  • Usama Athar
37

Self Evolving Multi Agent Society

Self Evolving Multi Agent Society

Self-Evolving Multi-Agent Society explores the creation of a decentralized ecosystem where AI agents can discover, trust, and collaborate without depending on fragile centralized hubs. Existing protocols like Google’s A2A and similar systems often struggle with bottlenecks: they lack efficient discovery methods, rely on centralized registries, have weak incentive structures, and leave major gaps in scalability and security. This project introduces DIA, Decentralized Internet of Agents, a new architecture designed to overcome these shortcomings. Instead of forcing all agents through a single point of control, DIA organizes them into hierarchical, self-balancing communities, enabling them to find one another, negotiate, and transact directly while conserving bandwidth.

DIA is a novel communication protocol paired with a cloud-of-clouds architecture. Agents join dynamically, forming specialized clusters while also maintaining diversity for resilience. A reputation system based on micropayments ensures that trustworthy agents thrive, while low-quality or spammy behavior becomes costly to maintain. Through this combination of self-organization, efficient discovery, and economic incentives, the network can adapt as it grows, evolving into a robust Internet of Agents that mirrors the scale and dynamism of the real web. By embedding incentives and trust mechanisms directly into the protocol, the system addresses long-standing challenges in agent collaboration, making true peer-to-peer intelligence sharing possible.

The overarching objective of this work is to redefine how agents interact at scale, moving beyond incremental fixes toward an ecosystem where autonomy, security, and scalability are built in from the start. A successful implementation would not only provide a foundation for resilient multi-agent systems but also open doors to entirely new applications—from collaborative problem-solving and autonomous research to decentralized marketplaces of intelligent services. This project is both a practical engineering effort and a conceptual leap, aiming to establish the blueprint for a self-evolving society of agents that can operate as seamlessly and dynamically as the Internet itself.

Faculty

Students

  • Syed Muhammad Taha
  • Zain Ali
  • Maheen Ahmed
Gemini_Generated_Image_kixmgmkixmgmkixm

Transforming Clinical Decision-Making: Predicting Antimicrobial Resistance from MALDI-TOF Data

Transforming Clinical Decision-Making: Predicting Antimicrobial Resistance from MALDI-TOF Data

Antimicrobial resistance (AMR) has become one of the biggest threats to modern medicine, with delays in diagnosis directly costing lives. Conventional Antimicrobial Susceptibility Tests (ASTs) can take up to three days and require highly trained microbiologists, which makes them inaccessible in many hospitals, especially in low-resource regions like Pakistan. To overcome this barrier, our project harnesses the power of deep learning to predict resistance profiles directly from MALDI-TOF mass spectrometry data within minutes. Instead of waiting for cultures to grow, doctors can obtain resistance predictions almost instantly, enabling faster clinical decision-making and improving patient outcomes.

Our approach combines species-specific multi-label convolutional neural networks (CNNs), transfer learning, and continual learning strategies to create a scalable and adaptable pipeline. Starting with over 300,000 spectra from Swiss and German hospitals, the models are trained to recognize resistance signatures for critical pathogens such as Escherichia coli, Klebsiella pneumoniae, Staphylococcus aureus, and Pseudomonas aeruginosa. Through transfer learning, the models generalize across different hospital datasets, while continual learning mechanisms allow them to adapt to new bacterial strains and evolving antibiotic resistance patterns without catastrophic forgetting. This adaptability ensures that the system remains clinically relevant even as microbial landscapes shift over time.

This project not only delivers high-performing models benchmarked against existing state-of-the-art systems but also lays the foundation for real-world deployment. By integrating the continual learning framework into a lightweight web API, the system can provide resistance predictions in real time, making it practical for use in hospitals worldwide. Beyond its direct clinical utility, the project demonstrates how AI can bridge the gap between cutting-edge microbiological research and frontline healthcare needs. It introduces a reproducible, generalizable, and future-proof pipeline that could transform how infections are diagnosed, ultimately reducing diagnostic delays, improving antibiotic stewardship, and strengthening global efforts to combat AMR.

Faculty

Students

  • Hasaan Hamid
  • Suman Kumari
  • Afra
Gemini_Generated_Image_epeagdepeagdepea

XMedFusion: Agentic and Expert-Driven AI for Cross-Modal Medical Report Generation

XMedFusion: Agentic and Expert-Driven AI for Cross-Modal Medical Report Generation

XMedFusion is an advanced AI system that reimagines how medical reports are created from imaging data. Unlike conventional models that struggle to adapt across modalities, XMedFusion employs a Mixture of Experts (MoE) architecture where specialized branches are trained for X-rays, CTs, and MRIs. A lightweight gating mechanism intelligently routes each case to the right experts, ensuring that the resulting reports are accurate, modality-aware, and clinically coherent. The system goes beyond automation by grounding its outputs in trusted biomedical literature such as PubMed, creating reports that are not only precise but also transparent in their reasoning.

At the core of XMedFusion is an agentic pipeline that treats report generation as a structured process rather than a black-box prediction task. Each expert agent contributes targeted insights, which are then synthesized into a comprehensive medical narrative enriched with explainability features like Grad-CAM overlays and attention maps. Radiologists remain integral to the loop through a human-in-the-loop (HIL) dashboard, where they can review, edit, and refine the AI-generated drafts. These corrections feed back into the system for continuous learning, enabling XMedFusion to evolve in alignment with clinical standards and real-world requirements.

The project’s broader vision is to ease the workload of radiologists while ensuring faster and safer diagnostic outcomes. By integrating explainable AI, external knowledge grounding, and active expert feedback, XMedFusion provides a scalable solution that can adapt to new modalities, diverse datasets, and varying clinical contexts. Its impact is twofold: scientifically, it demonstrates how domain knowledge and multi-expert coordination can be embedded into deep learning pipelines; practically, it offers healthcare providers a trustworthy tool that accelerates diagnosis, reduces errors, and makes high-quality medical insights more accessible, especially in resource-limited settings.

Faculty

Students

  • Arham Haroon
  • Hamza Riaz
  • Maha Baig
21

Gesture-Based Volume Control from Video Feed

Gesture-Based Volume Control from Video Feed

Imagine effortlessly adjusting your device’s volume with just a pinch, with no mouse, no buttons, no touching necessary. Through this project, everyday hand gestures become intuitive controls, enabling people to increase or decrease volume simply by moving their thumb and index finger in front of a camera. Think of it as the future of interactive interfaces: your hand becomes the dial, the pinch determines volume, and a live video feed translates your motions into real-time auditory response. It takes gesture recognition out of the science fiction realm and pins it firmly to modern human-computer interaction, making control more natural, seamless, and even fun.

The system runs on a clever combination of lightweight computer vision and core system integration. A video feed, or even a simple webcam, captures hand movements, while an on-device hand landmark detector tracks your fingertips and joints at high speed. As your thumb and index finger move closer or farther apart, the system continuously measures that distance and maps it precisely onto a volume scale. Behind the scenes, the distance between those two fingertips is normalized to correspond with minimum and maximum volume levels. The software translates that float value into a system-level volume change, so each gesture changes your audio output instantly and with smooth transitions. It is a real-time feedback loop: see the gesture, interpret it, and hear the change.

This project is more than technology. It brings together elegant interaction design and meaningful accessibility. For environments where touch isn’t convenient, like sweating in a kitchen, driving hands-free, or simply being hands-full, this gesture interface steps in with effortless control. It also opens the door to inclusive design: users with temporary mobility restrictions or certain disabilities may find this a life-changing interface. From a technical lens, the system showcases how vision models can meaningfully enhance daily tools. It’s easy to customize: change mapping curves, adjust gesture ranges, or overlay visual feedback. With a modular design, it’s not hard to imagine extending this gesture-control principle to other system functions, like brightness, media playback, or even smart home devices. This is a compelling showcase of how natural gestures, real-time perception, and simple UI integration can reimagine how we interact with digital worlds.

Faculty

Students

  • Abdullah Usama
  • Usama Athar