While stress-testing this site the way a recruiter would, I asked my own career assistant a question I had not prepared it for: whether I was open to remote work. The assistant did exactly what I had built it to do — it declined, because the answer was not in its evidence base. Correct behavior, wrong outcome. It is a question every recruiter asks, and the honest answer was a refusal.
Here is the part that matters: I knew about it within the day, because the assistant logs every question it cannot answer. The gap went into the evidence base, the assistant now answers that question with my actual availability, and a weekly test asserts it never regresses. That loop — measure, change, verify — is the most honest demonstration of how I work that I know how to give. So I built this entire site around it.
A portfolio that runs instead of describes
Twenty years of my applied work sits inside environments that cannot be shown. That is the standing problem for anyone who has built serious systems inside a defense enterprise: the better the work, the less of it you can put on a website. Most people solve this with adjectives. I decided to solve it with running systems.
The status page is the live monitoring surface of my real infrastructure — service health over service bindings, backup freshness with a 26-hour dead-man’s switch, self-healing activity counts. The demos run production methods on synthetic data in your browser: streaming anomaly detection, fuzzy entity resolution, a whole-life financial system with a Monte Carlo retirement engine. The assistant answers questions about my record from a structured evidence base, with guardrails, and is itself the strongest example of the retrieval-grounded systems I build. None of it is staged. If a backup goes stale, the status page says so, and I get paged.
The psychology is load-bearing
I am a licensed marriage and family therapist alongside the data science career, and this site is where the two halves stop being a novelty pairing and start being one discipline. Every interactive page has a small 🧠 toggle that annotates the design decisions in place: why the demos let you break things yourself (self-generated evidence beats presented evidence), why the sample data ships with flaws (a trend that only goes up reads as marketing), why the assistant is allowed to say “I don’t have that on record” (a system that never declines teaches people to over-trust it).
The newest demo makes that last point measurable. The Trust Calibration Lab asks you to review ten pieces of an intentionally imperfect AI’s work — first with bare, confident-sounding answers, then with calibrated confidence and the right to decline — and scores your own over-trust and under-trust in each condition. It is the guardrail philosophy from my production RAG systems, compressed into ninety seconds you perform on yourself. Your anonymous score joins the aggregate, which means the demo’s evidence gets stronger with every visitor.
Measuring in public
The status page now includes a section called “this site studies itself”: how many questions the assistant took this week, how many it refused, how many résumés were downloaded. Aggregate counts only — there is no individual tracking anywhere on this site, which is itself a design position. The interesting column is not the counts. It is what changed because of them. The remote-work gap above. A hedging habit in the assistant’s voice that is now banned in its instructions and tested weekly. A tendency to improvise answers about weaknesses, replaced with a written honest-limits statement.
There is also an A/B experiment running on the résumé page’s download button right now. I will publish the result either way, including if the boring variant wins, because an experiment you only report when it flatters you is not an experiment.
What stays private, and why that is part of the design
The autonomous system behind my job search — discovery, AI fit-scoring, tailored document generation — has a live console I do not publish, because it holds real pipeline data about real companies. The boundary is deliberate and disclosed. Twenty years in regulated environments taught me that what a system refuses to expose tells you as much about its builder as what it shows. This site tries to demonstrate both.
If you are evaluating me — as a recruiter, a hiring leader, an engineer, or someone’s screening agent — the invitation is the same: do not take my word for any of this. Measure yourself in the lab, interview the assistant, check the live numbers, and take the résumé with you. The site is the portfolio.
Leave a Reply