Sheet 01 — General arrangement · REV 2026.09

Agents that check their work. Systems that hold.

I am Edwin Knuth, a senior staff product engineer in Portland. I spent a decade on observability UI at New Relic, then two years owning the capture frontend of an FDA-cleared imaging product at Overjet. Now I build agents, and the part I care about is the difference between an answer that sounds right and one that is.

Open to senior through principal roles · Portland, OR · remote-friendly

Specification

Discipline
Agents and LLM systems · Product engineering · Observability
Grade
Senior / staff / principal
Availability
Immediate · remote
AGENTS · LLM SYSTEMS MCP SERVERS GROUNDING · EVALS OPENTELEMETRY FDA 510(K) REACT · TYPESCRIPT RAG · EMBEDDINGS ON-DEVICE AI PYTHON · SWIFT · KOTLIN OBSERVABILITY AGENTS · LLM SYSTEMS MCP SERVERS GROUNDING · EVALS OPENTELEMETRY FDA 510(K) REACT · TYPESCRIPT RAG · EMBEDDINGS ON-DEVICE AI PYTHON · SWIFT · KOTLIN OBSERVABILITY

Sheet 02 — Measured results

Numbers, taken off the drawing

Every figure below came out of shipped production work, not a slide. The context for each one is on the résumé sheet.

MEASURED IN PRODUCTION
510(k)

FDA cleared

K253930 — Iris imaging, capture frontend

99.5%

sessions finalized

capture completes through backend outages

1.4%

variance vs official

406,769 ballots read by a local vision model

0

memory cut

on-device ASR, Android peak RSS

Sheet 03 — Selected work

Things I built to find out if they were possible

These are personal builds, not employer work, which is why I can show you the source. The shipped product engineering is on the résumé sheet; this is what I do when I want to find out whether something is possible.

DWG 001

Clip Portal

Live

A clip search where every result opens on the frame it matched

A search over 300 public-domain films from the Prelinger collection, 66 hours, where each result is a frame at a second and clicking it seeks the player there. Pick clips into a set and an Evaluate page says what you would be licensing: runtime, decades, narrated or silent, near-duplicate pairs, a stratified sample, a manifest. The same search answers agents through three MCP tools, and a retrieval eval over 100 hand-written queries scores every path the site serves.

  • Next.js
  • TypeScript
  • Cloudflare Workers
  • Postgres + pgvector
  • Lance
  • SigLIP
  • MCP
  • OpenTelemetry
  • Python
frames embedded, 300 films
264,626 frames embedded, 300 films
recall at 10, 100 queries
0.39 recall at 10, 100 queries
scrub to seek, median
224 ms scrub to seek, median
DWG 002

Receipts

Open source

An investigation agent that has to show its receipts

An investigation agent for Honeycomb that works a scripted incident in a real Honeycomb environment over the hosted MCP and has to cite the query and rows behind every claim. An eval harness with injected faults and known ground truth grades it on whether the answer was right, not on which tools it called.

  • Python
  • Claude API
  • Honeycomb MCP
  • OpenTelemetry GenAI
  • pydantic
  • pytest
top hypothesis right, Sonnet
30 of 30 top hypothesis right, Sonnet
mean total score, -1 to 1
0.92 mean total score, -1 to 1
thirty investigations, Sonnet
$16.09 thirty investigations, Sonnet
DWG 003

Agent Blue

Client work

An operations agent that has to prove what it says

An operations platform I built alone. A conservation nonprofit in Alaska runs their day to day on it, along with a handful of other client orgs. Chat agent on one side, a drafting API on the other, and a fact checker between the output and the person reading it.

  • Python
  • TypeScript
  • Claude API
  • Cloudflare Workers
  • ChromaDB
  • Ollama
  • OpenTelemetry
chat platforms, one agent
3 chat platforms, one agent
test files
142 test files
commits since Jan 2026
1,400 commits since Jan 2026

Sheet 05 — Field notes

What the numbers said

Short notes from the two projects above. Each one is a claim, a table, and what the table does not say. Four minutes each.

  1. Sheet 05.7

    How it was built

    Two projects in eleven days, built through Claude Code with one loop, four steps, and a person who writes the issues and says merge. The counts from the session transcripts, what the review caught, and what the person did.

    4 min read

  2. Sheet 05.6

    The view you license from

    A clip search over 300 public-domain films where every result opens the player at the frame it matched. Pick clips into a set and one page says what you would be paying for. The page's numbers, the ones that are good and the two that are not, and a sampling thesis the eval turned down.

    3 min read

  3. Sheet 05.5

    Where the page is ahead of the tool

    The same clip search answers a person through a page and an agent through three MCP tools. One query run both ways, and a list of what the person gets that the agent cannot reach yet. The tools have a few things the page lacks too.

    4 min read

  4. Sheet 05.4

    The grader that charges for confidence

    Ten incidents scripted into a real Honeycomb environment, with the true cause in a file the agent never sees. A grader that scores the answer, and takes more off a wrong answer said with confidence than a wrong one said with doubt. Thirty of thirty right on one model. Six of thirty on another that passes the vendor's own process check.

    4 min read

  5. Sheet 05.3

    The budget that measured itself

    A trace of one request, stage by stage, is what explained a bad number. A search timeout had moved from eight seconds to fifteen on the strength of one measurement, and within ninety seconds of the deploy two searches were cancelled at fifteen point one.

    10 min read

Sheet 04 — Service record

Where the production scars came from

Mar 2024 — Aug 2026

Overjet

Senior Staff Software Engineer

Dental AI · Iris Intelligent Imaging System

  • Shipped the Iris capture frontend through FDA 510(k) clearance (K253930), the company’s first clearance in image capture and PACS.
  • Built offline-mode resilience so capture sessions finish through backend outages, driving a 99.5% session finalization rate. Converting the app to a PWA is what made it possible.
  • Scaled the capture product from one pilot clinic to thousands of clinics.

Jan 2022 — Feb 2024

New Relic

Staff Software Engineer

Observability platform · Logging

  • Shipped two LLM-powered features in New Relic AI for Logging, error analysis on log lines and automatic generation of parsing expressions, in production for enterprise customers in 2023.
  • Rearchitected data fetching across Logging UI components to meet performance SLOs and cut query timeouts on large datasets.
Full résumé Download PDF → Six years of New Relic APM and Browser work also on that sheet

Sheet 06 — Scope of work

What I am looking for

  • 01 Senior through principal, on a team where the hard problem is real. Title matters less to me than whether the problem is.
  • 02 Agents and LLM systems are what I build now. The part I care about is the gap between an answer that sounds right and one that is, and the structural work that closes it: citing the rows behind a claim, saying what went unchecked, writing evals that punish confident and wrong.
  • 03 Product engineering is the day job and always has been: the frontend of an FDA-cleared medical device, the UI of an observability platform used by enterprises. That is the half that turns a working prototype into something people rely on.
  • 04 Agent and AI tooling, developer tools, or observability. A decade of platform work sits under all three, and the on-device projects are where I learned what a model does when the budget is real.
  • 05 Remote, Portland-based. Comfortable being the person who owns the ugly reliability problem nobody has claimed.

[email protected]