INTRO · 0:00
Mark Wireman · 2026
The Self-Learning Multi-Agentic System
A deep technical walkthrough of an AI that researches, learns, validates, and persists mastery — completely autonomously.
INTRO · 0:30
The Long-Standing Challenge
Every AI session begins with amnesia.
No memory of past learning. No retained mastery. No compounding growth. Every run starts at zero.
INTRO · 1:00
The Answer
Learn. Remember. Master.
Learn
Six-phase loop, cycle after cycle
Remember
Semantic episodic memory across sessions
Master
Quantified threshold, validated by assessment
CONTROL TAB · 5:30
Architecture
Eight Specialized Agents. One Orchestrator.
01
Researcher
10-engine search + synthesis
02
Learner
Study materials + gap identification
03
Assessor
Generate 25–50 Q 4 difficulty levels
04
Evaluator
Answer · score update mastery
05
Implementer
Generate · validate repair code
06
Transcriber
yt-dlp + Whisper video ingestion
07
SkillGenerator
Package knowledge as SKILL.md
08
BaseAgent
Retry logic model fallback
No inter-agent imports. All state flows through KnowledgeState . The orchestrator is the only coordinator.
LEARNING LOOP · 14:30
The Core Architecture
The Six-Phase Learning Cycle
LEARNING
CYCLE
skip 1–2 if
all mastered
PHASE 1
Research
PHASE 2
Learning
PHASE 3
Assessment
PHASE 4
Self-Test
PHASE 5
Evaluation
PHASE 6
Gap ID
LEARNING LOOP · 17:00
Measurement, Not Assumption
Mastery Is Measured
Score Blending
Stability: one bad cycle can't erase accumulated mastery
LEARNING LOOP · 20:00
Gap-Directed Research
The Loop Is Genuinely Adaptive
PHASE 6
Gap ID
weak_topics[]
get_weak_topics
(threshold=0.85)
NEXT CYCLE
Phase 1
targeted queries
ONLY WEAK
SUB-TOPICS
identifies gaps
filters mastered topics
seeds gap-targeted queries
researched topics
This is not a pipeline. It is a feedback system — the outputs of each cycle are the directed inputs of the next.
MEMORY TAB · 36:00
Semantic Episodic Memory
Learning Transfers Across Sessions
PRIOR SESSION
Distributed Systems
mastery: 89%
EMBEDDING SPACE
cosine similarity retrieval
NEW SESSION
Microservices Patterns
seeds from prior context
Semantic retrieval — not keywords — automatically transfers relevant context. No manual tagging. No curation. The vector space is the memory.
IMPLEMENTATION TAB · 22:00
Validated, Not Just Generated
Code That Proves It Works
preflight
env · deps
syntax
parse · lint
contract
interface spec
assertions
invariants
config
settings · env
definitions
types · schemas
startup
launch · verify
test
unit · integration
install
pip · packages
normalize
sanitize output
smoke_test
basic probe
✓
The output is not a suggestion. It is a working artifact . That distinction matters when code will actually be deployed.
IMPLEMENTATION TAB · 26:00
Directed Repair, Not Random Retry
Failures Are Classified, Not Retried Blindly
VALIDATION
FAILURE
BUG TRIAGE
ML PIPELINE
classify · categorize · score
TARGETED
REMEDIATION
DIRECTED
RETRY
phase N exit code ≠ 0
external ML — independent analysis
category + severity + fix directive
more specific each attempt
The repair is directed, structured, and gets more specific with each attempt. It doesn't spin randomly — it converges .
SKILLS TAB · 39:00
The Output Protocol
Every Run Produces a Reusable Artifact
<topic>-skill.zip
├── SKILL.md ← main artifact
├── references/
│ ├── sources.md ← URLs + confidence
│ └── assessment_qa.md ← gap evidence
└── implementation/
├── src/
├── tests/
└── requirements.txt
FULL FORMAT
9 sections · complete depth Deep-dive assistant context
SUMMARY FORMAT
4 sections · token-optimized Limited context windows
Loadable as a slash command context by any AI agent in the ecosystem
SKILLS TAB · 41:30
Meta-Cognitive Architecture
The System Carries a Model of Itself.
ai_self_learning_system.md
→
Derived from primary literature: Voyager · STaR · AlphaEvolve Darwin Gödel Machine
A knowledge artifact about how self-learning systems work — available as context while the system is learning.
This is not a coincidence. It is intentional meta-cognitive architecture.
GROUNDBREAKING · 47:00
Eight Reasons This Is Different in Kind
01 Mastery is measured, not assumed — quantitative pass/fail, not qualitative summaries
02 Research is multi-modal — papers, repos, video lectures, all ingested uniformly
03 The loop is genuinely adaptive — outputs of cycle N are directed inputs of cycle N+1
04 Cross-session memory is semantic — transfer learning without fine-tuning or curation
05 Implementation is validated, not just generated — 11-phase working artifact
06 The repair loop uses external analysis — bug triage ML, directed remediation
07 Security is built into the architecture — MITRE ATLAS-mapped injection filter at ingestion
08 Outputs are reusable expertise artifacts — SKILL.md bootstraps future AI work on related topics
GROUNDBREAKING · 51:00
The Honest Limitation
An Amplifier, Not an Oracle.
The Limitation
Mastery is self-referential . The system creates the test, takes the test, and grades the test. Internal consistency is not external ground truth.
What It Is
An amplifier of expert judgment. It compresses days of research and debugging into hours — but the final "is this right for production?" is still a human question.
OUTRO · 54:00
A Knowledge Acquisition Engine for the AI-Native Era
Every run adds to institutional memory.
Every session produces a reusable artifact. Every artifact accelerates the next run.The system creates its own successors.
8 Agents · 6 Phases · 11-Phase Validation
10 Search Engines · Semantic Memory · SKILL.md