Real projects. Real decisions moved.
A selection of engagements delivered personally by Dr Ahmed Younes — some through AYD, others earlier in his career — across advocacy, disinformation, conversational AI and customer insight, for research institutions, governments and businesses.
LLM-as-a-Judge Evaluation Framework for Conversational AI
An automated evaluation pipeline for a customer-service voice AI — GPT generates adversarial scenarios, drives live conversations and judges outputs across design bugs, logic bugs and quality. Scaled from 7 to 42 test cases across 14 categories.
RAG Autocomplete System for Accessibility Audits
Built and deployed a RAG service that auto-drafts structured fields for accessibility audit issues — retrieving the most similar past incidents and drafting a grounded completion rather than inventing detail. Currently in beta ahead of full rollout.
NLP Analytics Pipeline for Unstructured Text
Built a portable topic-modelling and RAG analytics pipeline — 22,000+ records indexed in Qdrant, a multi-page Dash dashboard, and agentic thematic allocation making the data directly accessible to stakeholders.
Agentic Accessibility Scanner
An LLM tool-call layer over deterministic Playwright automation that detects WCAG violations and captures targeted screenshots — validated against real Wikipedia pages, catching a real detection bug along the way.
AI Strategy & Roadmap for a National Accessibility Platform
Led AI strategy and roadmap planning for a national government digital accessibility programme — a seven-layer platform architecture, a three-agent automation model, and direct contribution to bid and proposal development.
Pro-Kremlin Influence Network Mapping
Mapped pro-Kremlin influence networks across 7.8 million Telegram posts in France, Germany and Italy.
Global Fund Advocacy Evaluation
Evaluated social-media advocacy for the Global Fund's 7th replenishment across 5 markets and 43,000 posts.
Investor Sentiment in South African Energy
Built an investor-sentiment index for South Africa's energy sector, validated as predictive of investment flows.
Sweden's Global Image Under Pressure
Analysed the impact of two sensitive international incidents on Sweden's image across 9 countries and 7 languages.
Disinformation Landscape Mapping
Built a multilayered topic-modelling system to profile disinformation actors on social media.
China's Public Diplomacy on Twitter & Facebook
Mapped China's diplomatic social-media activity — 432,800 messages across 372 accounts, with 102,883 classified across 9 themes in 4 languages.
Customer Review Topic, Theme & Sentiment Analysis
Analysed ~13,000 customer reviews for a guided-walking-holiday operator — topic modelling, thematic annotation and a zero-shot correction layer lifted macro-F1 from 0.69–0.79 to 0.85–0.95.
LLM-as-a-Judge Evaluation Framework for Conversational AI
An automated evaluation pipeline for a customer-service voice AI — GPT generates adversarial scenarios, drives live conversations and judges outputs across design bugs, logic bugs and quality. Scaled from 7 to 42 test cases across 14 categories.
RAG Autocomplete System for Accessibility Audits
Built and deployed a RAG service that auto-drafts structured fields for accessibility audit issues — retrieving the most similar past incidents and drafting a grounded completion rather than inventing detail. Currently in beta ahead of full rollout.
NLP Analytics Pipeline for Unstructured Text
Built a portable topic-modelling and RAG analytics pipeline — 22,000+ records indexed in Qdrant, a multi-page Dash dashboard, and agentic thematic allocation making the data directly accessible to stakeholders.
Agentic Accessibility Scanner
An LLM tool-call layer over deterministic Playwright automation that detects WCAG violations and captures targeted screenshots — validated against real Wikipedia pages, catching a real detection bug along the way.
AI Strategy & Roadmap for a National Accessibility Platform
Led AI strategy and roadmap planning for a national government digital accessibility programme — a seven-layer platform architecture, a three-agent automation model, and direct contribution to bid and proposal development.
Pro-Kremlin Influence Network Mapping
Mapped pro-Kremlin influence networks across 7.8 million Telegram posts in France, Germany and Italy.
Global Fund Advocacy Evaluation
Evaluated social-media advocacy for the Global Fund's 7th replenishment across 5 markets and 43,000 posts.
Investor Sentiment in South African Energy
Built an investor-sentiment index for South Africa's energy sector, validated as predictive of investment flows.
Sweden's Global Image Under Pressure
Analysed the impact of two sensitive international incidents on Sweden's image across 9 countries and 7 languages.
Disinformation Landscape Mapping
Built a multilayered topic-modelling system to profile disinformation actors on social media.
China's Public Diplomacy on Twitter & Facebook
Mapped China's diplomatic social-media activity — 432,800 messages across 372 accounts, with 102,883 classified across 9 themes in 4 languages.
Customer Review Topic, Theme & Sentiment Analysis
Analysed ~13,000 customer reviews for a guided-walking-holiday operator — topic modelling, thematic annotation and a zero-shot correction layer lifted macro-F1 from 0.69–0.79 to 0.85–0.95.