Gian McCoy

Apps

Lead Enrichment app · AI Screening Pipeline · DEX Replay Engine

Three working AI applications I designed and run on my own infrastructure. They are not practice products. They are where I learn, every day, how AI systems behave in production: how they fail, what they cost, and how to keep them in check.

What this means for your practice

  • I have built Twilio voice and voicemail pipelines with transcription and call classification, so phone-side problems are familiar ground.
  • Rules handle the predictable calls and AI handles the rest. That keeps a receptionist cheaper to run and easier to check.
  • Every AI call in my own apps is logged and costed. That habit shapes what I look for when I review your receptionist’s call results with you each month.
See AI receptionist consulting →

Lead Enrichment app

Multi-channel prospect intelligence and CRM/voice automation: five capabilities in one production FastAPI application.

Currently operating across 2,656 California contractor prospects.

Architecture & Design

The Lead Enrichment app is a single FastAPI application consolidating five capabilities that would otherwise require five to seven separate SaaS subscriptions: a GoHighLevel CRM webhook router, a Twilio voice and voicemail pipeline with encrypted Matrix-protocol operator alerts, an SMB prospect intelligence engine with a multi-vendor enrichment waterfall, an SEO intelligence subsystem (GSC/GA4 sync, content gap analysis, RAG-grounded article briefs, NxN local rank tracker, GBP competitive audit), and a Microsoft 365 SMTP email warmup system orchestrated across 11 n8n workflows on ARM64 Docker.

The system's LLM abstraction layer routes calls across Anthropic Claude, OpenAI, Gemini, and local Ollama by tier (SIMPLE work routes to Claude Haiku 4.5, MEDIUM to Sonnet 4.6, COMPLEX to Opus 4.6), with a single environment variable switching providers globally. Cost controls are baked in: input-hash deduplication on scoring calls, deterministic temperature for classification, and structured per-call logging for cost observability across a single dashboard.

Gian designed the system and directed LLM tools (primarily Claude) to write the code. He owns the agent contracts, database schemas, grading rubric, abstraction seams, migration history, and operational discipline. The platform demonstrates what the architect-operator model looks like at single-tenant, production scale.

Key Features

  • Multi-provider LLM abstraction layer (Anthropic Claude / OpenAI / Gemini / Ollama) with tier-routed model selection
  • GoHighLevel CRM webhook integration with per-domain routers for Contact Created, Message Received, and Voicemail Transcript events
  • Twilio Voice + AMD voicemail campaigns; Twilio Lookup v2 phone-type classification; Matrix-protocol operator alerts with CRM context
  • Multi-vendor enrichment waterfall: BetterContact, Serper (5-pass), ZeroBounce, Twilio Lookup v2, Wappalyzer, Google Places, California CSLB (4-pass fuzzy match)
  • Self-hosted n8n on ARM64 Docker (11 workflows): email warmup tick, reply-tick, maintain, drip-board state transitions, GHL→Vikunja task push
  • pgvector RAG pipeline with 0.45 cosine threshold, IVFFlat index, and graceful fallback: anchors Claude article-brief generation to real career experience
  • Direct Microsoft 365 SMTP with warmup tick architecture: 15/day cap, 9 AM-5 PM PT window, Saturday 20% gate, 45-150 min reply gate
  • 200+ pytest tests with SSE-streamed test runner UI; 20+ specialized background workers with cancel-flag support

Tech Stack

FastAPIPythonPostgreSQLpgvectorAnthropic ClaudeOpenAIGeminiOllamaTwilioGoHighLeveln8nDockerMatrixReact 18ViteAWS EC2

AI Screening Pipeline

Rules-first screening pipeline with two-stage AI evaluation across five listing sources.

700+ listings processed and scored across five calibration profiles.

Architecture & Design

The AI Screening Pipeline is a two-repo system: a FastAPI backend that pulls listings from five sources (public APIs, an Apify scraper, and a custom Manifest v3 Chrome extension), and a React 18 + Vite + TanStack Query frontend with a Kanban-style triage board. Every listing is deduplicated by canonical key, keyword-matched against a reference bank, sorted into one of seven categories by a fully deterministic classifier, scored across four dimensions, and surfaced for a human to review.

Category assignment is rules-only by design. Hard gates exclude listings that are incomplete or out of scope before any keyword scoring runs. After the gates, weighted keyword scores and title bonuses decide the category. The two-stage AI step runs only after that: GPT-4.1-mini normalizes the raw listing into structured fields, then GPT-5 evaluates fit against a reference profile and past human decisions. Input hashes prevent paying twice for the same listing.

Claude Sonnet 4.6 generates tailored documents, all anchored to one canonical profile so every output draws on the same source of truth. A calibration loop compares the evaluator's calls with human overrides over time, so its thresholds keep improving. The same pattern (rules screen first, AI scores the rest, humans correct it) applies to screening calls, intake forms, or referrals in a practice.

Key Features

  • Deterministic classifier with hard-gate exclusions across 7 active categories
  • Source adapters for five listing sources, including an Apify scraper and a Chrome extension (Manifest v3)
  • Two-stage AI evaluation (GPT-4.1-mini structured normalization, then a GPT-5 fit decision with scores and dealbreakers)
  • Input-hash deduplication on GPT scoring calls: identical reruns become cache hits unless a rescore is forced
  • Calibration loop: evaluator thresholds tuned against the history of human overrides
  • Claude Sonnet 4.6 tailored document generation from one canonical profile
  • Weighted scoring across four dimensions per category
  • Kanban triage UI with category badges, AI-decision badges, and in-line stage transitions

Tech Stack

FastAPIPythonPostgreSQLOpenAI GPT-5GPT-4.1-miniAnthropic Claude Sonnet 4.6React 18ViteTanStack QueryChrome Extension (MV3)

DEX Replay Engine

Full-stack on-chain backtest workbench with hand-rolled indicators, Cartesian parameter sweeps, and AI-assisted reduction.

15 pytest files including integration tests; Cartesian sweeps over 10,000+ parameter combinations executable against in-memory cached bars.

Architecture & Design

The DEX Replay Engine is a full-stack Python + React workbench for analyzing and backtesting trading strategies against decentralized exchange data on EVM-compatible blockchains. It ingests on-chain swap events and block-level reserve snapshots into PostgreSQL, joins them into a canonical swaps + reserves relation, and aggregates to 5-minute and 15-minute USD-quoted OHLCV bars via materialized views. A React UI renders candlestick charts with indicator overlays, an equity-curve plot, and a timeline scrubber for replaying any block or timestamp.

Five hand-rolled strategies ship with the engine: RSI (Wilder smoothing), MACD, EMA Cross, Donchian Breakout, and ATR Trailing Stop. No TA-Lib dependency: every formula is visible and auditable. The backtest loop handles warmup pre-loading, close-vs-next-open execution mode, DEX fee deduction, stop-loss and take-profit thresholds, and the assembly of an equity curve, trade log, signal log, and indicator series. The parameter sweep loads bars once into memory, generates the Cartesian product of every strategy knob's range under constraint hooks, runs the backtest on each combination, and scores each run on a risk-adjusted objective.

The AI sweep reducer is stateless and bounded: it receives only the numerically pre-reduced result set (percentile statistics and a stable parameter region computed in code first) and asks OpenAI to summarize it in natural language. The model cannot hallucinate a parameter recommendation because the parameter region is fixed before the model sees it. This is the same deterministic-first pattern used across all three apps.

Key Features

  • Hand-rolled trading indicators with no TA-Lib dependency: RSI (Wilder smoothing), MACD, EMA Cross, Donchian Breakout, ATR Trailing Stop
  • Cartesian parameter sweeps with constraint hooks (e.g., buy_threshold < sell_threshold) and risk-adjusted scoring: final_equity − λ × max_drawdown − μ × total_fees
  • AI sweep reducer: numerically pre-reduces results (top-K region, percentile stats), then asks OpenAI for natural-language summary, so no parameters are hallucinated parameters
  • CPMM AMM swap math (dx = dy × x / (y + dy)) for historical price-impact reconstruction; UniV3 structure stubbed with HTTP 501
  • Materialized-view-backed OHLCV pipelines (rb_token_bars_mv, ui.pool_bars_usd_q_5m, ui.pool_bars_usd_q_15m) for sub-second queries
  • 15 pytest files including integration tests against real DB snapshots and unit tests using mock Bar objects
  • Pool Story transaction deep-dive: before/after reserve state, surrounding swap context, animated SwapFlowCenterpiece
  • Trading engine package (app/trading/) with zero FastAPI imports, fully unit-testable in isolation as a standalone library

Tech Stack

FastAPIPythonPostgreSQLSQLAlchemy 2.xOpenAIReact 18Vitereact-financial-chartsLeafletTanStack Query

One Operating Model Across All Three

Determinism First

Rules-based code runs before any LLM call. Hard-gate exclusions, SQL-derived grading, and keyword classifiers handle the obvious decisions. The model runs only on the residual where probabilistic reasoning earns its cost.

Provider-Agnostic LLM Layer

Agents declare a tier (SIMPLE / MEDIUM / COMPLEX), not a model string. The abstraction layer resolves to the right model (Haiku 4.5, Sonnet 4.6, Opus 4.6, GPT-4.1-mini, GPT-5), and a single env var switches providers globally.

Architect-Operator Model

Gian designed the systems and directed LLM tools (primarily Claude) to write the code. He owns the schemas, contracts, migration order, auth posture, abstraction seams, and operational discipline. The LLM is the implementer.

Want this kind of care applied to your practice’s phones and systems?