AI Evaluation Engineer
I design the benchmarks that make AI models smarter.
role: "AI Evaluation Engineer"
background: ["tech", "finance", "people"]
superpower: "Designing the scenarios
that break AI models —
then teaching them to win"
status: building|
Most AI engineers can build the system. Most annotators can label the data. I engineer the scenarios that teach models how to think.
I started in the trenches — diagnosing DSL line faults at 2 AM, scripting network configs, bouncing line cards on Adtran systems. I earned a BBA in Computer Information Systems — a degree that fused technology with finance, accounting, economics, and quantitative methods — then spent 5+ years as a Risk Advisor, where I didn't just sell policies — I systemized risk management. I built Salesforce workflows, automated outreach funnels, trained teams on data integrity, and turned analytics dashboards into revenue machines.
Finance runs deep. I manage risk for families and businesses, engineered lainsurancequotes.com as a digital brand to capture and qualify local lead traffic, and actively trade cryptocurrency using a custom framework rooted in market psychology, liquidity analysis, and capital allocation. Money management isn't a side skill — it's how I think.
Now I'm channeling all of it — the technical depth, the financial acumen, the people instincts — into AI evaluation and training data engineering. I design complex agentic tool-use scenarios and web-source reasoning tasks that generate RLHF/SFT training data for large language models. I work backwards from deterministic answers to construct multi-hop reasoning chains, author golden trajectories with atomic fact verification, build rubrics, and analyze model failures with surgical precision — all through rigorous QA pipelines with 10–15+ validation checkpoints per task.
The market is flooded with people who can use AI. It's starving for people who can design the benchmarks, engineer the training data, and diagnose the failures that make models actually improve. That's me.
Designed a highly complex, 6-tool agentic workflow to test an LLM's cross-tool entity matching and temporal reasoning. Engineered a synthetic, noise-heavy environment (800+ objects, 300+ arrays) requiring the model to filter decoys and track state across systems.
Utilized a 2-run methodology: Run 1 passes cleanly. Run 2 introduces strict developer instructions to force measurable, attributable failure.
Designed a natural-language question requiring an LLM to combine historical, geographical, and municipal-record clues across 5+ authoritative sources to reach a single numeric answer. Tests multi-hop reasoning, temporal anchoring, entity disambiguation, and document-based extraction.
Designed a cross-domain reasoning task requiring an LLM to traverse 8 distinct authoritative sources — spanning ancient history, geography, institutional affiliations, national identity, award recipients, journal metadata, and deep-learning model metrics — to extract a specific Testing MAE from a peer-reviewed PDF.
Authored rubric-based prompts against video clips engineered to stump LLMs through precise audio-visual integration. Each prompt required genuine multimodal reasoning — answers could not be confirmed from audio or video alone — bounded by a 1–2 minute context window defined by two observable events.
Prompt tracked fine-grained physical actions performed by Dr. Watson under hypnosis across a 66-second window [52:07–53:13], synchronized against Dr. Onslow's verbal commands.
Stellar AI
Handshake AI
Allstate Insurance — Paul Mims Agency
Self-Employed
Pattern recognition, data pipeline thinking, and building decision systems under uncertainty — core competencies for designing AI evaluation scenarios and training data.
Kim Duke State Farm
CRM architecture, workflow automation, team training on technical tools, and data-driven funnel optimization — directly transferable to AI evaluation design and rubric-based quality systems.
Family Care
Full-time care for elderly family member. Demonstrated empathy, adaptability, and problem-solving under pressure — the human-centered instincts that inform how I design technology for real people.
Kim Duke State Farm
Marketing automation, CRM segmentation, and analytical tool configuration — foundational logic for designing structured evaluation workflows and training data pipelines.
CenturyLink
Bayou Internet
University of Louisiana at Monroe
2011 — 2016
Hybrid foundation: finance, accounting, economics, marketing, statistics, systems analysis, database development, network design, network security, and project management.
Insurance Certification · License #LA-867875
Licensed across life, health, property, and casualty lines.