Skip to content
QA OS

Roadmap

A dependency graph, not a calendar.

No dates. Every arrow below is a real dependency — feedback loops are structurally meaningless before execution is real, and specialization is premature before evaluation exists to measure whether "specialized" is actually better.

  1. 1
  2. 2
  3. 3
  4. 4
  5. 5
  6. 6
  7. 7
  8. 8
  9. 9
  10. 10

Training / evolution plan

Where "training" actually belongs — answered directly, not hedged

1

Retrieval + structured memory first

What's already built — cheapest, most interpretable, easiest to debug when it's wrong, and it's already producing real signal today via the Module 1 ↔ Module 3 loop.

2

Context / prompt specialization second

Org-specific terminology, severity conventions, and known-risky-area weighting encoded as retrieved context, not baked into weights. Still fully interpretable, still cheap to iterate.

3

Evaluation-driven iteration third

The actual gate. Without packages/evaluation being real, there's no honest way to know if steps 1–2 made anything better — only that they changed something.

4

Fine-tuning / preference learning last, and narrow

Only once there's real labeled trace volume — hundreds of real approve/reject/correct decisions per task type — and only for specific sub-tasks where retrieval and context genuinely plateau. Full-system fine-tuning is explicitly the wrong first move: expensive, hard to debug, and it throws away the interpretability retrieval-based memory gives for free.

What's missing, concretely

Feedback should be human-labeled not just as approve/reject (already captured) — but why, in structured form: wrong risk level? Missing test type? Tone mismatch with team convention? That's currently missing, and it's a real, concrete next feature — not vague future work.

Partnership vision

The foundation here is independently built and honestly scoped. The next real step — organization-specific specialization, real historical data, real feedback at volume — isn't something one person builds alone in good faith; it needs a real QA environment's data, workflows, and judgment.

If that's an interesting problem to you or your team, I'm open to technical collaboration, pilot work, or figuring out what a deeper partnership could look like. This isn't a sales pitch — it's an open engineering direction with a clear next chapter.

Talk to me about this direction