INTRO · 0:00
Mark Wireman · 2026
The Self-Learning
Multi-Agentic System
A deep technical walkthrough of an AI that researches, learns, validates, and persists mastery — completely autonomously.
INTRO · 0:30
The Long-Standing Challenge
Every AI session
begins with amnesia.
No memory of past learning. No retained mastery.
No compounding growth. Every run starts at zero.
INTRO · 1:00
The Answer
Learn. Remember. Master.
Learn
Six-phase loop,
cycle after cycle
Remember
Semantic episodic
memory across sessions
Master
Quantified threshold,
validated by assessment
CONTROL TAB · 5:30
Architecture
Eight Specialized Agents.
One Orchestrator.
01
Researcher
10-engine search
+ synthesis
02
Learner
Study materials
+ gap identification
03
Assessor
Generate 25–50 Q
4 difficulty levels
04
Evaluator
Answer · score
update mastery
05
Implementer
Generate · validate
repair code
06
Transcriber
yt-dlp + Whisper
video ingestion
07
Skill­Generator
Package knowledge
as SKILL.md
08
Base­Agent
Retry logic
model fallback
No inter-agent imports. All state flows through KnowledgeState. The orchestrator is the only coordinator.
LEARNING LOOP · 14:30
The Core Architecture
The Six-Phase Learning Cycle
LEARNING CYCLE skip 1–2 if all mastered PHASE 1 Research PHASE 2 Learning PHASE 3 Assessment PHASE 4 Self-Test PHASE 5 Evaluation PHASE 6 Gap ID
LEARNING LOOP · 17:00
Measurement, Not Assumption
Mastery Is Measured
Score Blending
70%
New Cycle
+
30%
Prior History
Stability: one bad cycle can't
erase accumulated mastery
Difficulty Weights
Recall
1.0×
Comprehension
1.5×
Application
2.0×
Analysis
2.5×
LEARNING LOOP · 20:00
Gap-Directed Research
The Loop Is Genuinely Adaptive
PHASE 6 Gap ID weak_topics[] get_weak_topics (threshold=0.85) NEXT CYCLE Phase 1 targeted queries ONLY WEAK SUB-TOPICS identifies gaps filters mastered topics seeds gap-targeted queries researched topics
This is not a pipeline. It is a feedback system — the outputs of each cycle are the directed inputs of the next.
MEMORY TAB · 36:00
Semantic Episodic Memory
Learning Transfers Across Sessions
PRIOR SESSION Distributed Systems mastery: 89% EMBEDDING SPACE cosine similarity retrieval NEW SESSION Microservices Patterns seeds from prior context
Semantic retrieval — not keywords — automatically transfers relevant context.
No manual tagging. No curation. The vector space is the memory.
IMPLEMENTATION TAB · 22:00
Validated, Not Just Generated
Code That Proves It Works
preflight env · deps syntax parse · lint contract interface spec assertions invariants config settings · env definitions types · schemas startup launch · verify test unit · integration install pip · packages normalize sanitize output smoke_test basic probe ✓
The output is not a suggestion. It is a working artifact. That distinction matters when code will actually be deployed.
IMPLEMENTATION TAB · 26:00
Directed Repair, Not Random Retry
Failures Are Classified,
Not Retried Blindly
VALIDATION FAILURE BUG TRIAGE ML PIPELINE classify · categorize · score TARGETED REMEDIATION DIRECTED RETRY phase N exit code ≠ 0 external ML — independent analysis category + severity + fix directive more specific each attempt
The repair is directed, structured, and gets more specific with each attempt.
It doesn't spin randomly — it converges.
SKILLS TAB · 39:00
The Output Protocol
Every Run Produces
a Reusable Artifact
<topic>-skill.zip
├── SKILL.md ← main artifact
├── references/
│ ├── sources.md ← URLs + confidence
│ └── assessment_qa.md ← gap evidence
└── implementation/
├── src/
├── tests/
└── requirements.txt
FULL FORMAT
9 sections · complete depth
Deep-dive assistant context
SUMMARY FORMAT
4 sections · token-optimized
Limited context windows
Loadable as a slash command context by any AI agent in the ecosystem
SKILLS TAB · 41:30
Meta-Cognitive Architecture
The System Carries
a Model of Itself.
ai_self_learning_system.md
→
Derived from primary literature:
Voyager · STaR · AlphaEvolve
Darwin Gödel Machine
A knowledge artifact about how self-learning systems work —
available as context while the system is learning.
This is not a coincidence. It is intentional meta-cognitive architecture.
GROUNDBREAKING · 47:00
Eight Reasons This Is Different in Kind
01Mastery is measured, not assumed — quantitative pass/fail, not qualitative summaries
02Research is multi-modal — papers, repos, video lectures, all ingested uniformly
03The loop is genuinely adaptive — outputs of cycle N are directed inputs of cycle N+1
04Cross-session memory is semantic — transfer learning without fine-tuning or curation
05Implementation is validated, not just generated — 11-phase working artifact
06The repair loop uses external analysis — bug triage ML, directed remediation
07Security is built into the architecture — MITRE ATLAS-mapped injection filter at ingestion
08Outputs are reusable expertise artifacts — SKILL.md bootstraps future AI work on related topics
GROUNDBREAKING · 51:00
The Honest Limitation
An Amplifier,
Not an Oracle.
The Limitation
Mastery is self-referential. The system creates the test, takes the test, and grades the test. Internal consistency is not external ground truth.
What It Is
An amplifier of expert judgment. It compresses days of research and debugging into hours — but the final "is this right for production?" is still a human question.
OUTRO · 54:00
A Knowledge Acquisition Engine for the AI-Native Era
Every run adds to
institutional memory.
Every session produces a reusable artifact.
Every artifact accelerates the next run.
The system creates its own successors.
8 Agents · 6 Phases · 11-Phase Validation
10 Search Engines · Semantic Memory · SKILL.md
01 / 15
Title
← → SPACE · click · swipe  ·  N for notes