Mohammed Amro//Vancouver, BC//12+ yrs in ML
AI Engineering
Leader.
- Agentic AI in production
- Regulated clinical AI
- Teams built from zero
I turn organizations with no AI capability into ones that ship regulated AI on a repeatable platform.
- AI org built from zero
- 0→14
- AI org built from zero
- engineers & scientists
- projects delivered
- 17
- projects delivered
- across 5 areas
- Health Canada clearances
- 2
- Health Canada clearances
- + FDA 510(k) submission
- faster report turnaround
- 35%
- faster report turnaround
- eight-agent LLM system
- environments / countries
- 5 / 2
- environments / countries
- one CI/CD pipeline
01/About
Trustworthy AI, shipped by teams I built.
I've spent 12+ years in machine learning — first building models by hand, then building the teams and platforms that ship them. Most recently I led AI at a Series A medical-imaging company operating across the US and Canada, where I grew the AI organization from zero to 14 engineers and scientists and took it from research to regulatory clearance to live clinical workflow.
My work lives where AI has to be trustworthy, not just impressive: diagnostic models cleared by Health Canada, an eight-agent LLM system radiologists use to draft reports, and the evaluation discipline that makes "accurate" a claim that survives a regulator — facility-disjoint validation, confidence intervals, and error measured against how much radiologists disagree with each other.
I hold an M.Sc. in Data Science & Engineering and an MBA, and I care as much about the hiring pipeline, the release gate, and the board conversation as I do about the architecture.
01
Made “accurate” defensible
Facility-disjoint validation, bootstrapped confidence intervals, inter-reader baselines and imbalance-aware model selection became standard practice — not per-project choices.
02
Trusted nothing unmeasured
Proved DICOM tags systematically wrong, built ground truth from three-radiologist consensus, and benchmarked a well-marketed foundation model — then kept the in-house one.
03
Turned one-offs into platforms
One serving framework, 157 versioned annotation configurations, shared prompt tooling. Each new model cost less to ship than the last.
04
Built the capability, not just the output
Hired deliberately from Canadian master's and PhD graduates instead of competing for scarce senior talent — real depth, within a Series A budget.
Toolkit
Leadership
Team building 0→14 · Hiring & assessment · Distributed teams (3 time zones) · Delivery predictability · Board & exec communication
Agentic AI
Multi-agent systems · Orchestration & tool calling · MCP · LangGraph · LangChain · CrewAI · Retrieval of prior context · Prompt engineering
AI / ML
LLMs & generative AI · Computer vision · Segmentation & landmark detection · Transformers & CNNs · NLP
Evaluation & Responsible AI
LLM-as-judge · Hallucination measurement · Human-in-the-loop gates · Subgroup & fairness analysis · Bootstrapped CIs
Platform & MLOps
PyTorch · TensorFlow · MONAI · MLflow · DVC · CI/CD with automated V&V · GCP · AWS · Kubernetes · Event-driven microservices
Regulated AI
Health Canada clearance · FDA 510(k) · GMLP · SaMD lifecycle · HIPAA · DICOM PS3.15
02/Career Journey
From hands-on ML to building the organization.
Jun 2025 — Aug 2026
Vancouver, BC · US & Canada
07Executive Director, Product Management & AI Development
Series A medical-imaging company
- Architected an eight-agent LLM system for radiology report generation, shipped into clinical workflow across five environments — 35% faster report turnaround.
- Required generative output be measured before release: an evaluation harness with LLM-as-judge scoring, bootstrapped CIs and subgroup breakdowns — near-zero measured hallucination and zero false-positive recommendations on a 330-report stratified benchmark.
- Led a 15+ person cross-functional org across three time zones at 95%+ on-time milestone delivery; owned 14+ AI products end to end.
+ 2 more− Less
- Program manager for a $30M strategic enterprise partnership, with formal governance across product, engineering, regulatory and support.
- Established the Good Machine Learning Practice framework; participated in the FDA 510(k) submission for the XR Chest and XR Spine Finding Detectors.
May 2024 — May 2025
Vancouver, BC
06Director of AI Research & Development
Series A medical-imaging company
- Directed AI R&D across X-ray, CT, MRI and clinical NLP, leading 10+ researchers and engineers from ideation to cleared deployment.
- Cut manual annotation effort 50% with model-assisted labeling — including mining structured supervision from ~31,000 free-text radiology reports.
- Set the evaluation standard: facility-disjoint validation, confidence intervals, and error benchmarked against measured radiologist inter-reader variability.
+ 1 more− Less
- Commissioned a build-vs-adopt study of a medical foundation model; the in-house model won, preventing a costly platform commitment.
Aug 2022 — Apr 2024
Vancouver, BC
05Senior Machine Learning Manager
Series A medical-imaging company
- Built the AI organization from zero to 14 engineers and scientists — sourcing, assessment, offers and onboarding; new hires shipped production features within two months.
- Directed consolidation of clinical model serving onto one shared platform: seven model repositories, one CI/CD pipeline, five environments in two regulatory jurisdictions.
- Made automated verification & validation a release gate — releases blocked unless case matrices passed against live deployments.
+ 1 more− Less
- Cut training-to-deployment cycle time 45% through automated MLOps (MLflow, DVC) on GCP and AWS.
Dec 2021 — Jul 2022
Vancouver, BC
04Machine Learning Manager
Series A medical-imaging company
- Led cross-disciplinary ML teams delivering medical imaging and NLP systems.
- Evaluated and onboarded three vendor platforms, cutting external tooling costs 20%.
- Raised team NPS by 25 points across two performance cycles through a culture of rigor and code ownership.
Apr 2019 — Nov 2021
Vancouver, BC
03Senior Machine Learning Engineer
1QBit
- Architected and delivered XrAI, a Health Canada-cleared deep learning model for COVID-19 diagnosis from chest X-rays — 60% faster radiologist triage at peak pandemic demand.
- Authored an ML platform codebase later adapted as the foundation of my next organization's production ML platform.
- Built GAN- and CNN-based bone age estimation models for pediatric radiology (+18% diagnostic accuracy vs baseline reads).
Feb 2014 — Mar 2019
Doha, Qatar
02Data Scientist
Sidra Medicine
- Developed deep learning models for early-stage skin cancer detection — 94% sensitivity in clinical validation.
- Built ML queuing analytics that cut average wait times 22% across three outpatient departments.
- Integrated ML into EHR workflows with clinical teams, reducing manual data entry 35%.
Sep 2017 — Jun 2019
Doha, Qatar
Part-time · concurrent
01Teaching Assistant, Applied Deep Learning
Hamad Bin Khalifa University
- Lab sessions and one-on-one mentorship in deep learning, CNNs and TensorFlow for 40+ graduate students.
Education
2026
MBA, Business Administration
Edgewood University · Madison, WI
2018
M.Sc., Data Science & Engineering
Hamad Bin Khalifa University · Doha, Qatar
2019
Graduate Certificate, Data Science
Harvard Extension School · Cambridge, MA
2001
B.Sc., Computer Science
Alexandria University · Alexandria, Egypt
Certifications
03/Portfolio
17 projects. Five areas. One platform underneath.
Delivered under my leadership, ordered by visibility — though in practice the data foundation came first and the platform is what made the diagnostic work repeatable. Each links to a Case Study.
Area A—Regulated Diagnostic AI[5]
Models accurate enough for clinical measurement — the clearance path.
- A.01
MR Lumbar Spine Labeling & Abnormality Detection
36-landmark vertebral localization on sagittal MRI, accurate to within the range radiologists disagree with each other.
- mean landmark error
- 2.07 mm
- mean landmark error
- per-vertebra, facility-disjoint
- 96.2%
- per-vertebra, facility-disjoint
- A.02
XR Long Bone Labeling
Six bone segmenters in one self-routing container, and the migration that retired a legacy TensorFlow 1 pipeline.
- validation Dice, six segmenters
- 0.79–0.97
- validation Dice, six segmenters
- V&V cases per release
- 52 + 8
- V&V cases per release
- A.03
XR Chest View Position Classifier
A classifier that proved the DICOM headers were wrong — then replaced them.
- facility-disjoint accuracy (CI 0.887–0.935)
- 91.1%
- facility-disjoint accuracy (CI 0.887–0.935)
- mislabeled studies uncovered
- 148
- mislabeled studies uncovered
- A.04
XR Knee View Classifier & Abnormality Detector
When no labels existed, an LLM turned 31,000 narrative reports into training supervision.
- 5-class view accuracy
- 99.2%
- 5-class view accuracy
- reports mined into labels
- ~31k
- reports mined into labels
- A.05
MR Spine Sequence Classifier
A 2.5-D classifier that decides whether every downstream spine model gets to run.
- macro accuracy & F1
- 0.80→0.92
- macro accuracy & F1
- dataset scale-up
- 8.5×
- dataset scale-up
Area B—Production ML Platform[3]
Stopped rebuilding infrastructure per model; made each new one cheap to ship.
- B.01
Imaging Inference Platform
One shared framework so each new clinical model ships as a thin plug-in, not bespoke infrastructure.
- dissimilar AI services, one platform
- 4
- dissimilar AI services, one platform
- model repos, one CI/CD pipeline
- 7
- model repos, one CI/CD pipeline
- B.02
Series Organizer — MR Sequence Classification Service
300 series classified in under two seconds, from metadata alone — enforced in CI, not claimed.
- for 300 series, CI-enforced
- <2.0 s
- for 300 series, CI-enforced
- standardized sequence types
- 11
- standardized sequence types
- B.03
Series Registration
Aligning prior and current scans from three landmark pairs, ~4× faster than the library baseline.
- target registration error
- ~0.41 mm
- target registration error
- faster than volume registration
- ~4×
- faster than volume registration
Area C—Data Foundation[3]
Made the clinical archive legally usable, analyzable and annotatable.
- C.01
DICOM De-Identification Pipeline
Removing identity without destroying research value — the legal gate for every model.
- parallel, resumable
- 32-way
- parallel, resumable
- files provenance-stamped
- 100%
- files provenance-stamped
- C.02
Archive-Scale DICOM Metadata Extraction
Turning million-file binary archives into analysis-ready tables — the input to every cohort.
- files per archive
- ~1M
- files per archive
- curated tags, or all
- ~40
- curated tags, or all
- C.03
RedBrick Annotation ToolKit
One config and one command per annotation project — every operation audit-logged for regulators.
- versioned project configs
- 157
- versioned project configs
- operations audit-logged
- 100%
- operations audit-logged
Area D—Applied Generative AI[2]
An eight-agent clinical LLM suite, and the evaluation behind a build-vs-adopt call.
- D.01
Report Assistant — AI Radiology Reporting Suite
An eight-agent LLM system radiologists use to draft reports — measured, not asserted.
- faster report turnaround
- 35%
- faster report turnaround
- false-positive recommendations
- 0
- false-positive recommendations
- D.02
MedImageInsight Foundation-Model Evaluation
A well-hyped foundation model, evaluated properly on internal data, lost to what the team already had.
- AUC with a light adapter
- 0.62→0.86
- AUC with a light adapter
- in-house AUC — it won
- 0.868
- in-house AUC — it won
Area E—Clinical & Commercial Ops[4]
AI and data work extended into quality, finance and self-service analytics.
- E.01
Core Data Dashboard — Self-Service EDA
Bespoke SQL requests productized into a filterable tool non-engineers use themselves.
- tables in one reusable join
- 7
- tables in one reusable join
- row chunked ingestion
- 200k
- row chunked ingestion
- E.02
Radiologist Peer-Review Analysis Dashboard
A comparable quality score for peer review — with the full distribution always beside it.
- analysis views
- 5
- analysis views
- ratings → count-weighted GPA
- 1–5
- ratings → count-weighted GPA
- E.03
Radiologist & Facility Payment Engine
A monthly spreadsheet ordeal turned into one auditable click — that refuses to price what it can't.
- rate-card sheets, one pass
- 26
- rate-card sheets, one pass
- auditable monthly report
- 1-click
- auditable monthly report
- E.04
CPT Exam Codes Dashboard & Version Comparator
Thousand-row manual catalog diffs replaced by an automated change report.
- change types detected
- 4
- change types detected
- manual row-by-row diffs
- 0
- manual row-by-row diffs
04/Digital Twin
Ask my Digital Twin_
An AI version of me, grounded only in my verified career record. Ask how I built the team, how the eight-agent system works, or how I make AI accurate enough for a regulator.
- → Answers from my career record only
- → Won't invent numbers
- → Salary & availability: email me
Mohammed's Digital Twin
AI twin — may be imperfect
Hi — I'm Mohammed's AI twin, answering only from my verified career record. Ask me about my work, how I lead, or any project on this site.
Try asking