3 practices published · 16 candidates held · 2 withdrawn · 8 admitted companies · 0 work samples · updated 2026-08-27
senku. aifirst / practice
AI-first registry · practice

What the companies do, as exercises one person can perform

Oleg Malkov is choosing where to work next; this list is what the companies he grades are recorded doing, restated as exercises one person can perform.

Stage

3 practices published · 16 candidates held · 2 withdrawn · 8 admitted companies · 0 work samples.

How a practice gets here

A PRACTICE REACHES THIS PAGE THROUGH ADMITTED ROWS 8 COMPANIES ADMITTED 2 SUPPORTERS NEEDED 16 HELD 2 WITHDRAWN 3 PUBLISHED HERE a candidate is held until enough admitted rows attest it, and a blind reviewer confirms the action from the passages alone. no company is named here before its row is graded.
Drawn from the data: the live admitted count, the declared threshold, and how many practices each side of it holds.

These practices appear in the captured, hashed extracts of admitted companies' own engineering writing. Open the citations. Here is how one person performs each at personal scale, and here is one person's graded record of doing so.

What it does not claim: That performing them gets anyone hired. That is the hypothesis this page exists to test, and until the tests report, the page says so.

The five conditions a practice clears before it is published
  1. Supported by captured extracts from at least two ADMITTED registry rows, every cited claim current.
  2. Each extract shows the human action. A result figure or a named system on its own does not qualify.
  3. The blind extractor and reviewer split has run: the reviewer receives the passages without the proposed name and writes the action they show. A description that drops the observable kills the candidate, with no tie to break.
  4. The exercise names its shared observable, pointed at inside the supporting extracts.
  5. It files under the dimension whose evidence produced it.

2 is the design's declared default, and the first extraction pass kept it on 2026-08-24. Raising it to 3 would cut the candidate list from 21 to 10 and concentrate what survives on the three most-published companies, which would make this a description of who writes most.

The practices

State what the result was measured against

Oleg's observable: baseline

published 2026-08-27 · human decision · Measured production results

Exercise and evidence from 3 admitted companies

When you claim an improvement, name the control, the baseline, or the thing it displaced, in the same sentence as the number.

Checkable artifact: The published claim, carrying its comparison.

Evidence: GitHub, Ramp, Spotify.

Record the alternatives you rejected and the requirement that decided it

Oleg's observable: alternatives

published 2026-08-27 · human decision · Named operating systems

Exercise and evidence from 2 admitted companies

Before building something that already exists elsewhere, write down what you looked at, what you rejected, and the specific requirement that made you build.

Checkable artifact: The decision record naming each alternative and its reason for rejection.

Evidence: Spotify, Stripe.

Run the agent where colleagues can read the session

Oleg's observable: colleagues can read

published 2026-08-27 · human decision · Measured production results

Exercise and evidence from 2 admitted companies

Work with the agent in a shared, readable place, and keep the transcript.

Checkable artifact: The readable transcript of the session.

Evidence: Ramp, Shopify.

The weakest qualifier in the set. One supporter states it as a constraint and forbids the private case; the other is six words inside a sentence about adoption and says nothing about what running in public means operationally. The extraction reported its count both with and without this candidate.

The 16 held candidates and their grading record

The extraction record names every company, quotes every passage and carries every URL. It is private, and it stays private, because naming a company on this public surface is a claim about that company and this house makes that claim only through an admitted row. What each candidate carries here is the hash of the passage it rests on, which is the citation identifier and names nobody.

Gate the merge on a human review of agent-produced work

Do not merge an agent-authored change until a human review of it is recorded.

Checkable artifact: The pull request, showing a review event dated before the merge.

held · needs 2 admitted companies · 8 admitted today · Named operating systems · 5 captured passage(s)

Oleg's answer, held: Work going live without human review. Human review happens in production. Sdlc is fully executed by agents who play different roles including qa

Why held: The answer moves review into production, while this candidate requires a human review before merge.

At one person the reviewer is the author, so the exercise proves only that a review step happened. A second pair of eyes is the half one person cannot supply. The degradation is stated on the practice.

Attribute the agent in the record of the change

Name the agent as a co-author of the work it produced, in the commit or in the byline.

Checkable artifact: The commit trailer, or the author list on the published thing.

held · needs 2 admitted companies · 8 admitted today · Engineer authored cadence · 2 captured passage(s)

Human grading: not recorded.

One of its two supporters does not clear the admission bar, so this candidate falls below the threshold the moment the publication rule is applied. It is also the only candidate filed under its dimension.

Do the steps that must not vary in code the model never touches

In your agent pipeline, make the invariant steps plain code and leave the model only what is left.

Checkable artifact: The workflow definition, showing which steps call a model and which do not.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 4 captured passage(s)

Oleg's answer, held: Our agents are not running on any sdk or self crafted loops, they are proper Claude code or codex harnesses. So enforcement is important but only can be done by the means available to steer the harnes

Why held: The answer rejects the candidate's self-crafted-loop premise and its observable fails R16.

Publish a named limitation of the approach you are recommending

In the same document that promotes an approach, name a specific thing it does not do or has not been checked for.

Checkable artifact: The limitation sentence, in the same published document as the recommendation.

held · needs 2 admitted companies · 8 admitted today · Adverse evidence · 3 captured passage(s)

Oleg's answer, held: a per-item human step that stops scaling

Why held: The answer names the reviewers' per-item limitation. This candidate asks for the act of publishing a limitation.

The loosest grouping in the set. The three supporting limitations are of different kinds: a capacity bound, an evaluation that cannot diagnose, and a checker with no checks of its own. They are grouped because the human action is one action and it produces one artifact. Split by kind, each supporter stands alone and the candidate dies.

Enumerate your system's failure modes by name

List the ways your own system fails, individually, before anyone asks.

Checkable artifact: The published list, one named mode per line.

held · needs 2 admitted companies · 8 admitted today · Adverse evidence · 2 captured passage(s)

Oleg's answer, held: a written list of failure modes with their costs

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Give each task its own disposable full environment

Boot a complete working copy of the project per task and tear it down when the task ends, so the agent never works in your live checkout.

Checkable artifact: The boot and teardown definition committed in the repository: container file, worktree script, sandbox configuration.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 3 captured passage(s)

Oleg's answer, held: isolated environment per agent run

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Put the evidence a human would use to verify inside the agent's reach

Give the agent the logs, metrics, errors and rendered screens you would look at yourself, so it can check its own work.

Checkable artifact: The tool configuration granting those sources, plus a session transcript in which the agent cites one of them.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 3 captured passage(s)

Oleg's answer, held: the agent wired into the same instruments an engineer uses

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Write the knowledge the agent needs into the place it actually loads

Take what you know and the agent does not, write it into the file the agent loads, and stop relying on it to find the information elsewhere.

Checkable artifact: The committed context or instruction file, and the agent's run using it.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 2 captured passage(s)

Oleg's answer, held: written-up knowledge placed where the agent will read it

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Harvest what a human corrected back into the agent's instructions

When you correct the agent, write the correction into its instructions or into a check, so the same correction is never made twice.

Checkable artifact: The diff to the instruction file or the check, traceable to the correction that caused it.

held · needs 2 admitted companies · 8 admitted today · Human and societal consequences · 2 captured passage(s)

Oleg's answer, held: a human correction fed back into the agent's instructions or tooling

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Kept separate from p10 on the artifact test: p10 produces a context file, p11 produces a dated change to one whose provenance is the failure that prompted it. Merging them loses the observable that separates writing documentation from maintaining it.

Build a set of cases with known-correct answers and gate changes on it

Write cases whose right answer you already know, run the agent against them, and ship no context or configuration change that fails them.

Checkable artifact: The committed case set with its expected answers, and a results table per configuration.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 3 captured passage(s)

Oleg's answer, held: an eval suite of cases with known answers

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Make each run emit a machine-readable record of itself, and measure from that record

Have the agent's run write its own trace or usage record, and compute any number you publish from that record. An estimate does not count.

Checkable artifact: The stored run records, and a published figure traceable to them.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 3 captured passage(s)

Oleg's answer, held: traces and token usage

Why held: The submitted observable does not appear in both preserved blind descriptions, so R16 refuses publication.

Publish a production figure with its window, its population, and where it can be re-queried

Never publish a bare number about your own work. State the period it covers, the population it is drawn from, and the source someone could re-run.

Checkable artifact: The published figure carrying those three fields.

held · needs 2 admitted companies · 8 admitted today · Measured production results · 5 captured passage(s)

Oleg's answer, held: a usage number stated with its measurement window

Why held: The answer names the window but omits the candidate's population and re-query source, and it fails R16.

When you republish a figure, restate the earlier one beside it

Every time you update a number you have published before, put the previous value next to the new one, so the movement is visible without opening the old post.

Checkable artifact: The new publication carrying both values.

held · needs 2 admitted companies · 8 admitted today · Measured production results · 2 captured passage(s)

Oleg's answer, held: merged pull requests attributed to the agent

Why held: The answer names agent attribution. This candidate asks for the previous value beside the new value.

Cut personal data out of what the agent can reach, and say so

Decide what personal data the agent must never see, enforce it in the access path, and state the boundary publicly.

Checkable artifact: The access rule or exclusion in the configuration, and the published statement of it.

held · needs 2 admitted companies · 8 admitted today · Reproducible mechanism · 2 captured passage(s)

Oleg's answer, held: Ok

Why held: The answer records assent but supplies no observable, so R16 cannot publish it.

Keep the durable logic in your own harness so the model underneath can be swapped

Put the parts that outlive any model, the verification, the ordering and the recorded output, in code you own, behind a seam where the model is replaceable.

Checkable artifact: The harness, with the model named in one place, and a record of a swap that happened.

held · needs 2 admitted companies · 8 admitted today · Adverse evidence · 2 captured passage(s)

Oleg's answer, held: We keeping instructions and using flagship vendor harnesses

Why held: The answer chooses flagship vendor harnesses, while this candidate requires durable logic in a replaceable harness.

Express a repeated change as an automated pass that opens its own pull requests

When the same edit is needed in many places, write the pass that makes it and let it raise the pull requests, so you edit neither by hand nor file by file.

Checkable artifact: The transformation or scheduled job, and the pull requests it opened.

held · needs 2 admitted companies · 8 admitted today · Named operating systems · 2 captured passage(s)

Oleg's answer, held: pull requests

Why held: The phrase is common across unrelated candidates and does not resolve the reviewers' action mismatch.

The 2 withdrawn candidates

These answers were judgments on the candidate. They remain visible as rejected records and publish no company claim.

Publish the approach you built and then abandoned, with the observation that killed it

withdrawn after human grading · Adverse evidence

Oleg's answer, rejected: killed it

Why rejected: The submitted text is a rejection verdict without an observable.

Publish how much of the agent's output your own checks rejected

withdrawn after human grading · Adverse evidence

Oleg's answer, rejected: reject

Why rejected: The submitted text is a rejection verdict without an observable.

The same six dimensions, at the size of one person

WHAT A SELF-GRADED 2 HAS TO CITE 3 dated samples 30 days apart, end to end both numbers are counted at the push. a 2 that cannot show them is refused there.
Drawn from the data: the two numbers a self-graded 2 has to satisfy, both computed at the push.

Each dimension scores 0 absent, 1 partial or isolated, 2 repeated and concrete; maximum 12. A zero carries a review date and no digit, here as on a company row.

What each dimension asks of one person
Named operating systems
The systems you run and the job each one does, named in public: the harness, the gate, the queue. A 2 needs more than one of them, each with its job stated and an artifact a stranger can read.
Measured production results
Numbers from work that actually ran, each carrying the window it covers and the population it is drawn from. A 2 needs figures whose source a reader can re-derive from what you published.
Adverse evidence
What went wrong in your own work, published while it still costs you something: the approach you abandoned, the check that missed, the number that was wrong and what replaced it.
Reproducible mechanism
Enough of the sequence, the controls and the real failure output for a stranger to challenge or rebuild the mechanism without asking you anything.
Engineer authored cadence
A continuing dated record across different systems. One good piece of writing is a 1.
Human and societal consequences
Who can read this, whose data it touches, and what the system decides about a person, addressed as part of the design and published with it.
The rules every exercise runs under
  • Exercises run inside real work. A drill written to be graded is not evidence of anything.
  • The evidence is a published package: conditions, result, checks, the AI's share, limits, corrections. Every one of Oleg's repositories is private, so a link into a repository is a dead link to a stranger and the published package is the artifact.
  • Conditions are committed before implementation on every new sample. The compiled base is labelled retrospective where it is retrospective.
  • Every sample attributes assistance: what the person decided, what an AI produced, what another human contributed.
  • Where the artifact is code or a check, the sample carries the known-bad run that reaches the named assertion and fails. A guard is proven by the failure that motivated it.
  • Weekly cadence, with honest empty weeks. An empty week and a sample held back both read as not tested. Neither reads as failure.
  • The record's own zeros and ones name the next week's exercise. The record is what generates the curriculum.

The record

The same rubric, applied to the person who grades the companies.

Oleg Malkov

graded by a human · 2026-08-24

Named operating systems

No qualifying public evidence found in the reviewed sources · reviewed 2026-08-24

Measured production results

No qualifying public evidence found in the reviewed sources · reviewed 2026-08-24

Adverse evidence

No qualifying public evidence found in the reviewed sources · reviewed 2026-08-24

Reproducible mechanism

No qualifying public evidence found in the reviewed sources · reviewed 2026-08-24

Engineer authored cadence

No qualifying public evidence found in the reviewed sources · reviewed 2026-08-24

Human and societal consequences

No qualifying public evidence found in the reviewed sources · reviewed 2026-08-24

Why this record is here, and what a zero on it means

The person grading these companies is graded on the same six dimensions, from the same kind of public evidence, under the same gate. One record, two renderings: the method page shows it as the grader's disclosure, the practice page as the worked example.

Nothing has been compiled yet. The six dimensions stand at zero and each renders the dated line the companies get. A compiled base, when it exists, is labelled retrospective, because it is graded on work that was already done.

A zero here is bounded review absence with a date on it, exactly as it is on a company row. It is never a statement that the person cannot do the thing. A public record measures publishing practice. It cannot see private work, employer-owned work, or anything held back at Oleg's decision.

Work samples

0 work samples · the field contract is published first

An empty week reads as not tested. A sample held back at Oleg's decision reads as not tested. Neither reads as failure, and neither is quietly skipped.

Real work whose result can be published as an evidence package. Not work that lands in a public repository: every one of Oleg's repositories is private, and that definition would make every week empty by construction.

Building the same record yourself

The derivation rule, the six dimensions with their personal anchors, the repetition cap with its numbers, the sample field contract, and every check the gate applies. That is enough to run the same record against your own work. Nothing is ingested here, nobody is listed, and no person is graded by this house except its own editor.

Reuse terms: Not set yet. The method is on this page and on the grading method page.

Corrections

Corrections stay on the page with what was wrong. A company can challenge any claim or ask for removal. Write to oleg@mlkv.org.

References this design learned from

What was examined, and what was adopted, adapted or rejected
  • Open Badges 3.0, 1EdTech. ADOPTED: the evidence URL inside the claim, and the field separation of criteria, evidence, result, assessor, date and status. REJECTED: issuer-asserted credentials and self-issued badges. A badge adds an issuer's claim before any reader has established trust; this record is self-asserted, gate-checked, and re-gradeable by a stranger from the same evidence.
  • Kubernetes community membership ladder. ADAPTED: levels defined by counted, dated, public artifacts, which is the shape of the repetition cap and is now computed at the gate. REJECTED: sponsorship as the admission route, which needs a community to sponsor and consent to give.
  • GitHub Skills content model and exercise template. ADAPTED: work inside a repository, visible artifacts, automated feedback from repository events. REJECTED: course completion as hiring evidence. A workflow proves the condition it checks and nothing wider.
  • Exercism. ADAPTED: a stable exercise definition, iterations, and automated feedback separated from mentor feedback. REJECTED: synthetic isolated exercises as hiring evidence. Every exercise here runs inside real work that produces a real artifact.
  • SWE-bench and its dataset contract. ADAPTED: issue, base commit, patch, fail-to-pass and pass-to-pass tests, which are the replayable fields of a work sample and the known-bad discipline where the artifact is code. REJECTED: a benchmark score as a signal about a person.
  • OpenSSF Scorecard check requirements. ADOPTED in part: explicit current, stale, failed and inconclusive states, a known-bad proof for every check, and a correction route. Its per-check test requirement is the closest external precedent for running a gate's own selftest before the gate itself. REJECTED: the aggregate score, applied to people. No combined employability number exists anywhere in this design.
  • Julia Evans, the brag document. ADAPTED: the running, dated, self-maintained log as the record's maintenance form. REJECTED: self-attestation with no openable evidence. An entry carries something checkable or it does not enter.
  • roadmap.sh. REJECTED as a model: inclusion by editorial and community judgment is the drift the two-admitted-companies rule exists to prevent. It defines the adjacent what-to-learn category this page has to stay distinguishable from.
  • Learn In Public, swyx. ADAPTED: publishing the artifact is the second half of every exercise. REJECTED: no rubric and no evidence discipline, which is the whole difference between a habit and a record.
  • The GitHub contribution graph. REJECTED as evidence: volume is not mechanism. ADOPTED: the record lives on the person's own infrastructure with no central submission.
  • StaffEng, Will Larson. ADAPTED: derive the list from what practitioners actually do. REJECTED: interviews as the source. The source here is the captured, hashed extract.
  • sso.tax, the SSO Wall of Shame. ADAPTED into the exercise and artifact pair: a practice counts only when performing it leaves something a stranger can open. That is the project's checkable-fact discipline applied to a person.