Hi, I'm Emad.
I build intelligent systems.
I design and build end-to-end ML systems — from large-scale annotation pipelines and entity resolution to LLM-powered summarization and agentic AI products for compliance and due diligence workflows. Seven years of industry experience, primarily in regulated enterprise environments.
Experience
Download Resume PDFSenior Machine Learning Engineer II (Research)
Thomson Reuters
Toronto, Canada
Details & Impact
Applied ML at TR Labs for CLEAR, Thomson Reuters' due-diligence platform for KYC/AML screening and investigations, and the Checkpoint tax research platform. Built ML systems across the lifecycle — LLM features and agentic workflows, model-training pipelines, and secure data infrastructure — in a regulated environment.
NLP & Machine Learning Engineer
INAGO Inc.
Toronto, Canada
Details & Impact
Built NLP and language-understanding systems end to end — data preparation, model training, evaluation, and application integration — including transformer fine-tuning and research collaborations with university and industry partners.
Education
M.Sc. in Computer Science
York University
Toronto, Canada
GPA: 8.17 / 9
Thesis: Interactive Question Answering Using Frame-based Knowledge Representation
B.Sc. in Computer Engineering
Amirkabir University of Technology
Tehran, Iran
GPA: 17.18 / 20
Technical Skills
Projects
Entity Resolution Retraining & Snowflake Pipelines — CLEAR
2026Designed and built an on-demand model-retraining pipeline on SageMaker Pipelines, turning researchers' feature-generation, training, and evaluation code into reusable DAG steps. Enabled parallel experiments tracked in MLflow, cutting experiment turnaround by an estimated ~40%. Built proof-of-concept Snowflake processing and training pipelines and assessed the gaps for migrating AWS research workflows to Snowflake.
Agentic Investigation System — CLEAR Investigate
2025 - 2026Contributed to CLEAR Investigate, Thomson Reuters' commercially launched agentic AI product for investigations. Built a central configuration layer in PydanticAI for agent system prompts, tool definitions, and agent/subagent wiring, so the team could reconfigure investigation agents from one place instead of editing definitions across many files. Migrated the LLM-as-judge evaluation suite into a reusable repository prepared for CI/CD.
CLEAR Business AI — GenAI Report Summarization
2024Built the GenAI summary panel for CLEAR Business entity reports, live in production processing hundreds of reports daily. Designed selective XML extraction feeding targeted per-section LLM calls (business overview, key personnel, adverse media, liens and lawsuits), plus a Bing Search + LLM verification step that filters false matches between same-named entities, with each claim linked to its source in the report. Refined prompts and extraction logic through SME annotation rounds against product-manager exemplars.
Entity Resolution Data Infrastructure — CLEAR KYC/AML
2022 - 2024Built the research data infrastructure for CLEAR's entity resolution system (~800M entities, billions of documents, 700 idents/second). Designed a PII-isolated AWS research environment with SageMaker Studio, CloudFormation, and per-researcher profiles; defined a versioned schema reconciling two incompatible annotation sources, with Spark and Python pipelines to merge records and surface label conflicts; and automated the manual workflow into scheduled, repeatable SageMaker Pipelines.
Semantic Search Improvement — Checkpoint Tax Research
2021Shipped a sentence-embedding query-intent classifier for Checkpoint, Thomson Reuters' tax research platform, separating broad conceptual queries from specific tax lookups to promote higher-level documents — improving access to conceptual documents by 95% within a ~100ms latency budget. Identified the retrieval gap from SME and customer feedback.
Automated Question Generation from Documents
2020Fine-tuned T5 to generate questions from documents, cutting manual data-curation effort by 40%. Experimented with input representations and evaluated generated text with BLEURT in a research collaboration with university partners.
Domain-Specific Language Understanding
2019Trained domain-specific Word2Vec embeddings and LSTM-based NLU models, with interpretability testing to make model decisions transparent.
Conversational Question Answering
2018Built a conversational question-answering system using syntactic and semantic document analysis and automatic ontology generation, in a research collaboration with an industry partner.
Publications
Question-worthy Sentence Selection for Question Generation
Interactive Question Answering Using Frame-based Knowledge Representation
Time Aware Topic-based Recommender System
A Study on Prediction of User's Tendency Toward Purchases in Websites based on Behavior Models