About Focus Projects Contact
$ whoami

George Jieh

AI/ML Engineer & LLM Evaluation Specialist

LLM evaluation and RLHF training for frontier models, production AI pipelines, and multi-agent orchestration. Former Scale AI Oracle Tier trainer. Bay Area based, open to AI/ML developer and research roles.

~ rlhf: 12 training projects, Oracle Tier + Quality Analyst at Scale AI
~ stack: Python, PyTorch, scikit-learn, H2O AutoML, OpenAI API, Flask_
01.

About

AI/ML engineer and LLM evaluation specialist in the Bay Area. Nine years in financial services under FINRA and SEC regulatory oversight before pivoting to data science and machine learning. The finance career taught systems thinking under regulatory pressure. The AI work is where that thinking goes when the constraints are open-ended.

At Scale AI / Outlier (August 2024 - April 2025), earned Oracle Tier and Quality Analyst designation across twelve RLHF training projects spanning process-supervised code reasoning, multi-turn instruction following, multimodal visual grounding, tool-use accuracy, and adversarial prompt design. Worked across data science, Python engineering, generalist reasoning, and finance domain tracks. Reviewed and corrected peer annotations, built evaluation rubrics from scratch, and caught reward hacking in quality review. This work plausibly contributed to improvements across frontier models from multiple labs including OpenAI, Anthropic, Google DeepMind, Meta, and Microsoft.

Data science foundation from BrainStation (Sep 2023 - May 2024) and Google Data Analytics Professional Certificate. Won first place in the Google-sponsored capstone hackathon (Google Companion, April 2024), building a geospatial pedestrian safety scoring model from 1.9M+ Chicago crime records. Microsoft Certified: Azure AI Engineer Associate and Azure Data Scientist Associate. PCEP certified in Python fundamentals.

Currently rebuilding past projects to production quality while developing new AI systems. Operates agentic orchestration across multiple platforms including OpenAI Codex, Hermes Agent, and previously OpenClaw, for multi-agent coding workflows, knowledge management, and automated operations.

FINRA licensed (Series 7, 66, SIE). WSET Level 2 wine certified. Bilingual English and Mandarin. UC Berkeley BA in American Studies.

02.

Focus Areas

LLM Evaluation & RLHF

Multi-dimensional rubric-based evaluation, step-level process supervision, adversarial prompt design, reward hacking detection, and structured feedback writing for frontier model training pipelines.

RLHF Chain-of-Thought DPO Oracle Tier

Machine Learning

Supervised and unsupervised learning, classification on imbalanced datasets, geospatial modeling, AutoML frameworks, and systematic architecture evaluation with documented trade-offs.

scikit-learn PyTorch H2O AutoML Keras

Production AI Pipelines

LLM inference pipelines, dual-prompt architecture, structured output enforcement, production reliability safeguards, and RAM monitoring with auto-shutdown.

Python OpenAI API Ollama Flask

Agentic Orchestration

Multi-agent coding workflows and knowledge management across OpenAI Codex, Hermes Agent, and OpenClaw. SOUL.MD system prompt engineering, scheduled jobs, and cross-agent delegation.

OpenAI Codex Hermes Agent OpenClaw Ollama Cloud

Data Engineering

Web scraping with anti-bot techniques, SQLite databases with schema migrations, ETL pipelines, and data cleaning with documented rationale for every decision.

SQL Selenium Pandas NumPy

Finance Domain

FINRA-licensed (Series 7, 66, SIE) with 9 years in wealth management operations, portfolio analysis, and regulatory compliance. Ground-truth domain expertise for AI evaluation work in financial contexts.

Series 7/66/SIE Portfolio Analysis FINRA/SEC
03.

Projects

In Progress

FreshRSS Financial News Analyzer

Production LLM inference pipeline that ingests 24-hour financial news RSS feeds, parses articles, and routes inference through OpenAI API or local Ollama models based on availability.

Dual-prompt architecture with constant behavioral system prompt and dynamic per-article user prompt enforcing structured financial analysis output (sentiment extraction, ticker identification, market implications). Multilingual support, production reliability safeguards including RAM monitoring with 80% auto-shutdown, and reduced daily research time from 60+ minutes to approximately 5 minutes. Currently being rebuilt to production quality.

PythonOpenAI APIOllamaLLMNLPFinancial Analysis
In Progress

Wine Grape Climate Suitability Model

End-to-end ML pipeline predicting wine grape varietal suitability from regional climate data. Processed 143,000+ wine records and 60+ years of weather data (1961-2016).

Built custom Selenium scraper to supplement Kaggle datasets with Wine Enthusiast ratings. Engineered seasonal climate features aligned with vine growth cycles (budburst, flowering, veraison, harvest). Achieved 56% accuracy using H2O AutoML Distributed Random Forest across 19-class imbalanced dataset. Deployed as Flask web application with location-based prediction interface. Currently rebuilding with cleaner data sources, modular architecture, and Gradio for improved production readiness and accuracy.

PythonH2O AutoMLscikit-learnFlaskSeleniumSMOTE
04.

Contact

Open to AI/ML developer and research roles. Available for freelance AI training, evaluation, and annotation work on weekends.

contact@georgejieh.dev