Ongoing · Medium Priority

AI Careers

A working framework and evidence archive for proving judgment-layer work in a world where AI can fake the appearance of competence.

The problem

Senior practitioners and executives have a career evidence problem that gets worse the higher they sit. Execution-layer work leaves artifacts — code, screens, shipped features. Judgment-layer work doesn’t: the portfolio bet you argued for against the room, the org redesign that kept a team from wasting a quarter, the call you made with incomplete information that turned out right. None of that produces anything you can hand to an interviewer or drop into a promotion packet. I started this project after running into the sharpest version of the problem directly: a reader laid off after a quarter of genuinely holding a team together through uncertainty, with nothing to show for it because the work was entirely judgment.

AI made this worse, not better. A polished memo used to correlate with real understanding, because producing one took understanding. Now AI generates the memo in minutes with zero comprehension behind it. The signal that evaluators, hiring managers, and promotion committees have relied on for years — does the output look competent — has been severed from the thing it was supposed to measure. This isn’t a personal-branding problem. It’s an organizational one too: companies making layoff and promotion decisions off old signals are increasingly deciding on noise, while the people who actually understand the work have no recognized way to prove it.

The approach

This is a content research project, not a software tool — it’s the working archive and synthesis layer behind a series of articles on cogniflow-ai.com about the career-evidence problem. It runs on CogitOS, my personal research/content pipeline, and what it actually contains is a manifest of curated source material — around twenty Substack articles and YouTube interviews pulled into individual bundles (article + summary + extracted ideas per source) — plus a running ideas.md file where each extracted claim is tied back to its exact source passage, and a project brief (manifest.yaml) that states the argument the whole thing is building toward.

The core deliverable the research converges on is a concrete four-question framework for building portable evidence of judgment: situation (what was the context), decision (what call was made and why), risk (what could have gone wrong), change (what actually happened as a result). It’s explicitly designed to be sanitizable — strip the confidential specifics and what’s left is a reasoning record that proves comprehension without leaking anything proprietary. The point of doing this through the four questions as you work, not reconstructed from memory during a job search, is that the specificity that makes it credible is the first thing that decays.

The archive also tracks the adjacent argument threads that feed the main one: the T/C/L/D audit for sorting your own week into theatre, commodifying, on-the-line, and durable work; the “explanation artifact” as the commit message for judgment (what is this, why this approach, what would break, what did I learn); and the shift from credentials to a living transaction history of verifiable comprehension. Each idea in the archive is a load-bearing claim with a citation back to its source bundle, not a paraphrase — so the eventual articles are argued from evidence, not vibes.

What I learned

The clearest thing that came out of pulling twenty-plus sources together was how consistently the same structural mechanism shows up across completely different framings: AI didn’t erase the correlation between effort and competence, it erased the correlation between polished output and competence. Once I saw that pattern repeated — in the credentials-vs-artifacts argument, in the junior-hiring-collapse argument, in the T/C/L/D audit — the four-question framework stopped being one article’s suggestion and became the load-bearing idea the whole project should be built around, which is not where I expected to land when I started collecting sources.

Where this could go

The natural next step is turning the framework into something people actually use, not just read about — a lightweight template or prompt set for running the four questions on a real decision as it happens, sanitized for public use. At team scale, the same evidence gap shows up as an organizational blind spot: companies that cut junior roles for AI are now discovering (Klarna is the clean public case) that they eliminated the pipeline that used to produce senior judgment, and they have no record of the judgment the people they let go actually had. A version of this framework built into how a team documents decisions — not for compliance, but as a live “why we chose this” ledger — would give an org the same kind of legible, sanitizable evidence trail it’s currently asking individual employees to build for themselves, and it would make promotion and layoff decisions less dependent on who happens to be good at self-narration after the fact.

Takeaway

The evidence problem AI created for individual careers is the same evidence problem it created for organizations — nobody has a reliable record of who actually understood the work, and building one deliberately beats reconstructing one under pressure.

Text summarized and optimized using Anthropic’s models and reviewed by a human.