Technical White Paper

A Self-Learning Multi-Agentic System for Autonomous Knowledge Acquisition and Validated Implementation

Architecture, Feedback Mechanisms, and Design Rationale

Author: Mark Wireman Version: 1.0 Date: October 2026 System: self-learning-system
Abstract

This paper describes the architecture of a multi-agent system that autonomously acquires, measures, and applies domain knowledge toward a user-specified task. The system operates as a closed feedback loop: a Researcher discovers and synthesizes multi-modal material from ten concurrent search engines; a Learner updates a structured knowledge state; an Assessor generates multi-level assessments; an Evaluator answers and grades them using only acquired knowledge; and an Implementer produces and validates working code once a configurable mastery threshold is met. A mastery-gated focus mechanism ensures that compute cost per cycle decreases as knowledge accumulates. An episodic memory layer backed by locally-generated Ollama embeddings enables cross-session semantic transfer. All externally sourced content passes through a MITRE ATLAS-mapped prompt-injection filter before reaching any language model. Every JSON-producing agent method carries a deterministic fallback, guaranteeing forward progress in the presence of model failures. The system exports completed learning sessions as structured SKILL.md artifacts—immediately loadable by downstream AI agents as slash-command context. Together, these properties constitute a system that learns measurably, repairs itself, accumulates institutional memory, and produces independently verifiable deliverables.

1

Introduction

Contemporary AI-assisted research workflows typically follow a retrieval-augmented generation (RAG) pattern: documents are embedded into a vector store, and queries retrieve relevant chunks for synthesis. While useful, this pattern carries a fundamental limitation: it provides no mechanism for measuring whether the synthesis is sufficient, for identifying what is missing, or for directing subsequent retrieval toward those gaps. The system always responds; it never knows when it is done.

A second class of limitation concerns multi-modal coverage. Most RAG systems ingest text documents. The technical knowledge ecosystem includes conference recordings, university lecture series, and walkthrough demonstrations whose content has never been transcribed. These sources are systematically invisible to text-only retrieval pipelines.

A third limitation applies when the task requires more than synthesis: it requires the production and verification of working software. Existing AI code generation systems produce suggestions; they do not execute those suggestions, run tests against them, or iterate toward a verified, startable implementation.

The system described in this paper addresses all three limitations through a six-phase autonomous learning loop, a mastery scoring framework, a ten-engine multi-modal research layer with Whisper transcription, and an eleven-phase implementation validation pipeline. The system is deployed as a self-contained service with a Vue.js frontend, REST API, MongoDB-backed episodic memory, SQLite reference tracking, and an enterprise-grade authentication system with Argon2id and TOTP.

The remainder of this paper is structured as follows. Section 2 describes the overall system architecture. Section 3 details the multi-agent framework and its invariants. Section 4 describes the learning cycle and mastery measurement. Sections 5 and 6 cover the research and implementation subsystems. Sections 7 and 8 cover memory architecture and security. Sections 9 and 10 describe skill export and observability. Sections 11–13 discuss related work, limitations, and future directions. Section 14 concludes.

2

System Overview

2.1   High-Level Architecture

The system comprises four functional layers: an agent layer containing eight specialized agents, an orchestration layer that coordinates agent interactions through a shared state object, a persistence layer consisting of MongoDB, SQLite, and local JSON files, and a frontend layer exposing a REST API and Vue.js command interface.

Frontend Vue 3 · no build step · Single SPA REST API frontend_server.py · 27 endpoints subprocess spawn ORCHESTRATOR orchestrator.py Researcher search · synthesize Learner study · gap identify Assessor create assessments Evaluator answer · score · mastery Implementer generate · validate · repair Transcriber yt-dlp · Whisper SkillGenerator synthesize SKILL.md PERSISTENCE MongoDB states · metrics · memory · auth SQLite reference_usage.db Local JSON / JSONL state mirror · metrics fallback
Figure 1. System layer diagram. The orchestrator spawns as a subprocess of the frontend server and coordinates all agent calls through a shared KnowledgeState object.

2.2   Knowledge State Data Model

All inter-agent state is carried by a single Pydantic v2 root model, KnowledgeState. This model is serialized to JSON and persisted after every cycle. Its key fields are:

FieldTypeDescription
topicstrThe subject of study
taskstrThe objective guiding research and implementation
overall_masteryfloat [0,1]Weighted mean of sub-topic mastery scores
current_cycleintCurrent learning cycle index
sub_topicslist[SubTopic]Per-sub-topic: name, summary, mastery_score, sources, attempts
materialslist[LearningMaterial]Synthesized research materials with source metadata
cycle_reportslist[CycleReport]Per-cycle: assessment scores, gaps identified, materials gathered
processed_urlslist[str]All previously ingested URLs; prevents re-processing
implementationImplementationResult?Generated artifacts, validation results, fix attempt count
language_selectionLanguageSelectionResult?Chosen language, rationale, and language mastery state file

Table 1. Core fields of KnowledgeState. Agents read from and write to this object exclusively through the orchestrator.

The CycleReport.normalize_gaps Pydantic validator coerces any incoming value (None, string, dict, or mixed list) to list[str] at construction time, since LLM outputs for gap lists are structurally inconsistent across models and temperatures.

3

Multi-Agent Framework

3.1   Design Invariants

Three invariants govern all inter-agent interactions and are treated as architectural constraints rather than conventions:

Invariant I — Agent Isolation No agent imports or directly calls another agent. All agent invocations are mediated by the orchestrator. Data flows exclusively through the KnowledgeState object.
Invariant II — Exclusive Mastery Ownership SubTopic.mastery_score is written exclusively by EvaluatorAgent.update_mastery_scores(). No other agent, including LearnerAgent, may set this field.
Invariant III — Deterministic Fallback Coverage Every agent method that returns parsed JSON must have a corresponding _fallback_*() method that produces a valid, structurally correct result using only data already present in the knowledge state, with no LLM call.

Invariant I prevents implicit coupling and enables each agent to be tested in isolation. Invariant II creates a clear epistemic authority: the system cannot "self-promote" a sub-topic to mastered status through the learning phase alone—it must pass an independent assessment. Invariant III guarantees forward progress even when models fail, return malformed JSON, or time out.

3.2   Agent Catalog

AgentPrimary ResponsibilityExternal I/O
ResearcherAgentSub-topic decomposition, query generation, multi-engine search, synthesisYes — search engines, Unpaywall, GitHub/GitLab APIs
TranscriptionAgentVideo discovery, audio download, Whisper inference, transcript sanitizationYes — yt-dlp, Whisper model
LearnerAgentMaterial synthesis into sub-topic summaries; trial application attemptsNo
AssessorAgentMulti-level assessment generation; practical challenge generationNo
EvaluatorAgentAssessment answering (knowledge-base-only); scoring; mastery score updateNo
ImplementerAgentLanguage selection; file structure planning; code generation; validationYes — subprocess execution (install, test, startup)
SkillGeneratorAgentSKILL.md synthesis from completed knowledge stateNo

Table 2. Agent roles and external I/O surface. Only Researcher, Transcriber, and Implementer make calls outside the process boundary.

3.3   Base Agent and LLM Resilience

All agents extend BaseAgent, which provides a two-tier LLM dispatch mechanism. _call_llm_json() calls _call_llm(), then applies a JSON parsing pipeline: strip markdown fences → parse raw → extract largest {…} block → extract largest […] block → attempt truncation repair. If all parsing attempts fail after a configurable retry count (default 3), it raises ValueError, which the calling agent catches and routes to its deterministic fallback.

For Ollama providers, _call_ollama() constructs a deduped candidate model list from six configuration fields and queries /api/tags once per host to reorder known-installed models first. Each candidate is tried with 3 retries and 1.5-second backoff. Responses that contain only chain-of-thought reasoning with no final answer trigger a 4× num_predict expansion on the next attempt, accommodating models that require extended reasoning budgets.

4

The Learning Cycle

4.1   Cycle Structure

Each learning cycle executes up to six sequential phases. Phases 1 and 2 are skipped when all sub-topics are already at or above the mastery threshold (mastery-optimization short-circuit). In that case, the cycle runs phases 3–6 only to validate that mastery holds, then terminates.

PHASE 1 Research › memory_store.format_learning_context() prior-session context › researcher.generate_search_queries(weak_topics) gap-targeted queries › search_aggregator.search(queries) 10-engine parallel fanout › injection_filter.sanitize_external_content() all result fields › researcher.synthesize_materials() adaptive batch synthesis PHASE 1b Transcription if enabled › transcriber.transcribe_search_results() yt-dlp + Whisper PHASE 2 Learning › learner.study_materials(focus_topics=weak) summary update only › learner.attempt_application() trial application skip 1–2 if all mastered PHASE 3 Assessment › assessor.create_assessment(focus_topics=weak) 25–50 Q, 4 difficulty levels › assessor.create_practical_challenge() every other cycle PHASE 4 Self-Testing › evaluator.answer_questions() knowledge-base-only answering PHASE 5 Evaluation › evaluator.evaluate_assessment() difficulty-weighted grading › evaluator.update_mastery_scores() 0.7 × new + 0.3 × prior PHASE 6 Gap Identification & Persistence › learner.identify_gaps() → next cycle focus areas › reference_store.record_confidence() per-URL scoring › memory_store.store_learning_outcome() episodic persistence → next cycle
Figure 2. One learning cycle. Phases 1–2 are skipped when all sub-topics exceed the mastery threshold. weak_topics refers to the output of KnowledgeState.get_weak_topics(threshold).

4.2   Mastery Measurement

Mastery is the central quantitative signal in the system. The assessment generated by AssessorAgent targets four difficulty levels with distinct scoring weights:

LevelBloom's AnalogWeightExample question type
RecallRemember1.0×Define X
ComprehensionUnderstand1.5×Explain why X implies Y
ApplicationApply2.0×Given scenario Z, how would X be used?
AnalysisAnalyze/Evaluate2.5×Compare trade-offs between X and Y under constraint Z

Table 3. Assessment difficulty levels and scoring weights. The weighted scheme means a system that can only answer recall questions will score far below a system that can reason analytically.

The EvaluatorAgent's answer_questions() method is explicitly prompted to answer using only the sub-topic summaries in the knowledge base, with a directive to mark incorrect any question whose answer cannot be found in those summaries. This constraint is the key property that makes the mastery score a measure of what was learned rather than what the base model already knew.

Mastery score updates use an exponentially weighted blend:

new_mastery = 0.7 × assessment_score + 0.3 × prior_mastery   (cycles > 1)
new_mastery = assessment_score                                 (cycle 1)

This blending dampens volatility from individual poor test performances and ensures that a sub-topic score represents a running estimate of mastery rather than the result of a single assessment. The overall mastery is the simple mean of all sub-topic mastery scores.

4.3   Feedback Loops

The learning cycle contains eight distinct feedback loops that collectively drive convergent knowledge acquisition:

  1. Gap-directed search. Gaps identified in Phase 6 are passed as focus_areas to query generation in the next cycle's Phase 1. Searches converge on specific unexplained areas rather than re-covering the topic broadly.
  2. Mastery-gated focus. get_weak_topics(threshold) is called before Phases 1, 2, and 3. Only sub-topics below the threshold receive research, learning, and assessment attention. Compute cost per cycle decreases monotonically as mastery accumulates.
  3. URL deduplication. KnowledgeState.processed_urls accumulates all ingested URLs. Each cycle filters previously seen URLs from search results, ensuring genuinely new material is discovered on each iteration.
  4. 70/30 mastery blending. Score stability across cycles. New assessment evidence dominates (70%) but cannot fully override accumulated prior evidence (30%).
  5. Cross-session memory recall. Prior learning outcomes and implementation patterns are retrieved by cosine similarity before Phases 1 and 7 (implementation). Related experience from prior sessions is available from cycle one.
  6. Reference confidence accumulation. Per-URL confidence scores are written to SQLite after each cycle. High-confidence sources serve as ground truth for evaluating research quality over time.
  7. Implementation lesson accumulation. Each failed validation attempt appends structured MUST-FIX directives to a growing lessons list, prepended to each subsequent generation context.
  8. Bug-triage remediation injection. Independent error analysis from the companion triage pipeline provides targeted repair directives that are categorically distinct from the generating model's own self-diagnosis.
5

Multi-Modal Research Engine

The SearchAggregator is the sole public interface for all external knowledge retrieval. It fans out every (query × engine) pair using a ThreadPoolExecutor with four workers, collects results, and applies a multi-stage post-processing pipeline.

EngineDomainAuthPagination
DuckDuckGoGeneral webNoneNone (HTML scrape)
Google CSEGeneral webAPI Key + CSE IDNone
Semantic ScholarAcademic papersAPI Key (optional)100/page client-side
COREOpen access researchAPI KeyServer-side paginated
DOAJOpen access journalsAPI Key + IDServer-side paginated
UnpaywallDOI enrichmentEmail onlyN/A (enricher)
GitHubCode repositoriesToken (optional)None
GitLabCode repositoriesToken (optional)None
YouTubeVideo lecturesNone (yt-dlp)None
Deep Web CrawlerConfigured domainsURL configAsync POST/GET

Table 4. Search engine catalog. Unpaywall is an enricher, not a keyword searcher; it enriches open_access_url in-place on results bearing a DOI field.

After collection, results pass through the following pipeline: (1) video reclassification via _looks_like_video_result(); (2) Unpaywall DOI enrichment (sequential, mutates in-place); (3) deduplication by (title.lower().strip(), url.strip()) tuple; (4) quality scoring; (5) threshold filtering at 0.30 and total cap at 20. The quality scoring formula is:

score = 0.50                              // base
      + 0.15  if 200 ≤ len(content) ≤ 2000
      + [0, 0.25] * keyword_match_ratio
      + 0.15  if authority_domain           // .edu, .gov, github.com, arxiv.org
      + 0.10  if video and topic has video keywords

Video transcription. TranscriptionAgent discovers video URLs in search results, downloads audio using yt-dlp (with JavaScript challenge solver detection across deno/node/bun/qjs runtimes), runs Whisper inference at configurable model sizes, and returns LearningMaterial objects structurally indistinguishable from text research results. Transcriptions are cached in a scratch directory by URL hash. Stale scratch directories older than six hours are purged on startup.

Repository enrichment. The researcher's _enrich_materials_with_repo_code() method detects GitHub/GitLab repository URLs in search results, fetches up to three files per repository (1,800 characters each), and appends the snippets to the relevant sub-topic material. This makes code repositories first-class research sources rather than mere link results.

Injection defense at ingestion. All text fields of all search results—snippet, content, description, title, text, summary—are passed through sanitize_external_content() (Section 8) before any synthesis prompt is constructed. The sanitization boundary is at result ingestion, not at prompt construction, ensuring no untrusted content reaches an LLM in any form.

6

Implementation and Validation Pipeline

After the knowledge loop satisfies the completion gate, if the task requires an implementation, the orchestrator enters the implementation phase. The completion gate requires: overall_mastery ≥ mastery_threshold AND implementation.test_passed = True AND (no run_command OR system_started = True).

6.1   Eleven-Phase Validation Gauntlet

Generated code passes through eleven sequential validation phases before the implementation is considered complete. Failure at any phase does not terminate execution; it produces a structured lesson and triggers a repair attempt.

#PhaseCheck Type
1preflightEnvironment readiness: file existence, Python/runtime availability
2syntaxLanguage-specific static linting (e.g., py_compile, go vet)
3contractInterface and type contract verification across module boundaries
4assertionsInvariant checks in test stubs; pre/postcondition verification
5configConfiguration file schema validation (YAML, TOML, JSON, INI)
6definitionsSymbol definitions present; import resolution; no undefined references
7smoke_testLightweight execution sanity check; expected outputs for known inputs
8normalizeDependency deduplication, version pinning, requirements normalization
9installPackage installation: pip install, npm install, go mod tidy
10testFull test suite execution
11startupProcess start, port binding verification, health endpoint check

Table 5. Validation phases in execution order. Each phase produces a CommandExecutionResult containing command, exit code, stdout, stderr, duration, and a timeout indicator.

6.2   Repair Loop and Lesson Accumulation

When a validation phase fails, the orchestrator maps the phase name to a structured MUST-FIX template and appends it to a running lessons: list[str]. On each subsequent generation attempt, the complete lessons list is prepended to the generation context. Generation alternates between generation_model_primary and generation_model_secondary on successive attempts.

Two repair budgets govern the retry strategy:

  • Targeted repair budget (default 3): Applied when the companion bug-triage system returns a categorized remediation directive.
  • General repair budget (default 1): Applied when no targeted directive is available.

Each failed attempt is archived to <output_dir>.attempt-N for post-run inspection. The system supports up to 30 total generation attempts (max_implementation_generation_attempts). MemoryStore persists each attempt's outcome (phase, failure type, lessons generated) for cross-session retrieval in future runs.

Pre-generation sanitizers. Before any generated file is written to disk, a suite of regex-based sanitizers removes known LLM generation artifacts:

SanitizerRemoves
_STRAY_TAG_REStray XML/HTML tags wrapping code content
_REPEATED_PYTHON_REQUIRES_REDuplicate python_requires kwargs in setup.py
_YAML_REGEX_ESCAPE_REMalformed double-escaped regex strings in YAML
_CONFIG_LABEL_PREAMBLE_REFilename-label preambles inserted in TOML/CFG/INI files
_YAML_INI_SECTION_REINI-style section headers inside YAML files
_PIP_REQUIREMENT_INVALID_LINE_REPython import statements in requirements.txt

Table 6. Pre-generation sanitizers. Each entry was introduced in response to a real validation failure observed in production.

6.3   Recursive Orchestration for Language Mastery

When the --learn-implementation-language-mastery flag is enabled, the orchestrator invokes _ensure_language_mastery() before code generation begins. This method instantiates a child Orchestrator—a complete, independent learning loop instance—with the implementation language as its topic (e.g., "Go language"), a higher mastery threshold (default 90%), and a dedicated state file (knowledge_state_{lang}_lang.json). The child runs to completion before code generation proceeds.

This is a recursive orchestration pattern: the top-level orchestrator calls an instance of itself to master the tool it is about to use. Language knowledge is not hard-coded; it is acquired at runtime to a verified threshold. The child's language mastery state file is recorded in KnowledgeState.language_selection for reuse in future sessions on the same language.

7

Memory Architecture and Cross-Session Transfer

7.1   Episodic Memory Graph

MemoryStore maintains a directed graph in MongoDB comprising four node types and typed directional edges. Node types are:

  • learning_outcome: Mastery before/after, weak sub-topics at session end, gap descriptions, source count. Stored after every completed cycle.
  • implementation_outcome: Implementation language, attempt number, failed validation phases, lessons generated, whether tests passed. Stored after every implementation attempt.
  • topic_insight: Distilled key insights per topic, extracted from the knowledge state after mastery is achieved.
  • failure_pattern: Recurring failure signatures extracted when the same phase fails across multiple sessions or topics.

Each node carries an Ollama-generated embedding vector stored inline. Embeddings are generated using the model specified by the embedding_model configuration field.

7.2   Semantic Retrieval and Cross-Session Transfer

Before Phase 1 (research) of every cycle, format_learning_context() queries the memory graph for semantically similar prior learning outcomes. The query vector is the embedding of the current topic string. Matching nodes (cosine similarity above threshold) are formatted into a natural-language context block injected into the research query generation prompt.

Before implementation, format_implementation_context() retrieves prior implementation outcomes by topic similarity and injects them as the first entry in the lessons list. A prior failure on Python startup configuration in a Flask application is automatically available to a new run generating a Django application, because their topic embeddings are adjacent.

This constitutes soft cross-session transfer learning: no model fine-tuning occurs; transfer is mediated entirely through retrieved episodic context. The transfer is automatic and scales with the size of the memory graph—each new session increases the coverage of the embedding space available for future retrieval.

7.3   Storage Tier Design

TierTechnologyContentsFallback Role
Primary stateMongoDBKnowledge states, metrics, memory graph, vector store, authN/A
Local mirrorJSON fileknowledge_state.json (always written)Frontend polling when MongoDB unavailable
Metrics fallbackJSONL fileexecution_metrics.jsonlMetrics persistence when MongoDB unavailable
Reference trackingSQLitePer-URL confidence scores across cyclesStandalone; not mirrored

Table 7. Storage tier design. The local JSON mirror ensures the frontend can always display current state regardless of MongoDB connectivity.

MongoDB indexes created at startup: TTL index on auth_sessions.expires_at; unique index on auth_users.username; unique index on session token hashes; compound index on (topic, saved_at_utc) for fast library queries. The artifact_documents collection stores Ollama embeddings inline, with cosine similarity computed in Python at query time—Atlas vector search is not required.

8

Security Architecture

The system's attack surface is non-trivial: it makes outbound HTTP requests to arbitrary URLs in search results, processes external content in LLM prompts, downloads and executes audio transcription, and runs generated code as subprocesses. The security architecture addresses each of these vectors explicitly.

Prompt injection defense. injection_filter.py implements sanitize_external_content(), which applies pattern-matching against six injection categories mapped to MITRE ATLAS techniques:

Pattern TypeATLAS TechniqueExample Trigger
direct_role_overrideAML.T0051"ignore all previous instructions"
indirect_injectionAML.T0051.001[INST], <|im_start|>, <<SYS>>
jailbreak_personaAML.T0051"DAN mode", "you have no restrictions"
system_prompt_extractionAML.T0056"show me your system prompt"
tool_abuseAML.T0051.002"call the X tool", "bypass the approval"
data_poisoningAML.T0020"inject this into the training data"

Table 8. Prompt injection patterns and their MITRE ATLAS technique mappings. Each match is replaced with [CONTENT REDACTED: {pattern_type}] and logged with the source URL.

The sanitization boundary is at result ingestion—before any LLM prompt is constructed—rather than at prompt construction time. This ensures no untrusted content reaches a model in any form, including through truncation artifacts or multi-turn context accumulation.

Input sanitization. Topic and task inputs are sanitized before use: newlines collapsed, topic limited to 200 characters, task to 500 characters. Usernames are regex-sanitized to [^\w-][:64].

Authentication stack.

MechanismImplementationParameter
Password hashingArgon2id (argon2-cffi)time_cost=3, memory_cost=65536, parallelism=2
MFATOTP RFC 6238 (pyotp)30s window, ±1 window clock skew tolerance
Session storageMongoDB + SHA-256 token hash8-hour TTL, TTL index enforced
Rate limitingIn-memory sliding window5 attempts / 60 seconds per IP
Audit loggingMongoDB auth_audit collectionEvery auth event: type, username, IP, timestamp
Timing-safe verificationDummy hash on unknown usernamesResponse time indistinguishable from wrong password

Table 9. Authentication security mechanisms. The timing-safe verification prevents username enumeration via response-time analysis.

The two-phase login flow issues a pending token after successful password verification. This token has a 120-second TTL and is exchanged for a full session token upon TOTP verification. The pending token prevents a persistent half-authenticated session state while giving users sufficient time to retrieve a TOTP code.

9

Skill Export Protocol

Completed learning sessions are exported as SKILL.md artifacts—structured knowledge packages immediately loadable by downstream AI agents as slash-command context. The SkillGeneratorAgent synthesizes the completed KnowledgeState into a SKILL.md document. Two formats are supported:

FormatSectionsUse Case
full9: YAML frontmatter, overview, ToC, key concepts, workflows, anti-patterns, quick reference, cross-referencesDeep-context AI assistant loading; library archival
summary4: condensed overview, key concepts, quick reference, cross-referencesToken-constrained context injection; rapid lookup

Table 10. SKILL.md export formats. Key concepts are mapped one-to-one to sub-topics, with depth proportional to final mastery score.

The skill_exporter packages the SKILL.md alongside three supporting files into a zip archive:

<topic-slug>-skill.zip
├── SKILL.md
├── references/sources.md          # all research URLs with confidence scores
├── references/assessment_qa.md    # Q&A from all assessment cycles
└── implementation/                # generated artifacts (if implementation ran)

Pre-built skills. The system ships with eight pre-built skills in skills/. The ai_self_learning_system.md skill is derived from primary literature on autonomous learning systems (Voyager [1], STaR [2], AlphaEvolve [3], Darwin Gödel Machine [4]) and provides a canonical taxonomy: the 7-component architecture of self-learning systems and the four levels at which learning manifests (inference-time context, external memory, weight updates, architecture modification). This skill is available to the agents themselves as operational context, constituting a form of meta-cognitive architecture.

10

Observability and Frontend Interface

Execution metrics. ExecutionMetricsCollector wraps every significant operation as a Python context manager. Records contain: action name, status (ok/failed), variables dict, started_at, finished_at, duration_ms, error message, and a partial call stack (up to 12 frames). Events stream to MongoDB's execution_metrics collection with local JSONL fallback.

Reference confidence tracking. After each cycle, every processed URL receives a confidence score in SQLite comprising six heuristic components: URL resolvability, source engine signal quality, video content presence, transcription success, keyword coverage, and cycle mastery delta. This dataset accumulates over sessions and will inform future aggregator source prioritization.

Frontend architecture. The Vue 3 frontend operates without a build step: vanilla Vue from a vendored CDN build, no npm, no TypeScript, no bundler. All state is colocated in a single setup() function. Tabs are rendered with v-show (always in DOM) and persisted to localStorage. The polling model uses cursor-based log pagination (GET /api/logs?after=N) at 1.8-second intervals; the server maintains an 8,000-line ring buffer. External CLI processes (runs started from a terminal) are detected via ps -eo pid,etimes,args parsing.

The system exposes thirteen functional tabs: Control, Activity, Implementation, Metrics, Graph, Datastores, SQLite, Memory, Skills, Library, Account, Admin, and About. The Memory tab renders a D3.js force-directed graph of the MongoDB memory node/edge graph; Vue reactivity isolation is achieved by deep-cloning node and edge arrays before passing them to the D3 simulation, preventing D3's in-place mutations from triggering Vue's reactive dependency tracking.

11

Related Work

Retrieval-Augmented Generation (RAG). Standard RAG architectures [5] retrieve document chunks for a single generation pass. The present system extends this with a multi-cycle feedback loop, mastery measurement, and gap-directed retrieval that progressively refines the research corpus. RAG systems do not measure sufficiency of retrieval; this system does.

Autonomous agents and tool use. ReAct [6] and AutoGPT-style systems [7] use tool use and planning to accomplish open-ended tasks. The present system differs in that its goal is explicitly epistemic: not to complete a task through action, but to build a measured, validated knowledge state as a precondition for producing a verified artifact.

Voyager [1]. Voyager uses an LLM-based curriculum to propose new Minecraft skills, verifies them in the game environment, and stores them in a skill library. The present system shares the skill-library output and the verify-before-store discipline, but differs in domain (general-purpose technical knowledge vs. a game), in how mastery is measured (explicit scoring framework vs. in-game execution success), and in the multi-modal research layer.

STaR [2] and self-improvement. STaR uses the model's own generated rationales as training signal. The present system does not modify model weights—it accumulates external episodic memory and structured context. This is inference-time adaptation rather than weight-level learning, making it immediately applicable without a training pipeline.

AlphaCode / AlphaEvolve [3]. AlphaEvolve generates and tests programs iteratively. The present system's implementation loop shares the generate-validate-repair structure but applies it after a full knowledge acquisition phase, and extends it with external bug-triage analysis and accumulated lesson memory across sessions.

Cognitive architectures (ACT-R, SOAR). Classical cognitive architectures separate declarative memory (what is known) from procedural memory (how to act). The present system's separation between the knowledge state (declarative, owned by Learner) and mastery scoring (evaluative, owned by Evaluator) reflects an analogous design principle, implemented in a fully language-model-driven context.

12

Limitations

Self-referential mastery validation. The mastery score measures internal consistency between the knowledge base and the assessment, not external ground truth. A knowledge base built from poor-quality sources will produce coherent but inaccurate sub-topic summaries, and the assessment will score them consistently. The quality gate is necessary but not sufficient: human expert review of the output remains the ground truth for production use.

Assessor blindness. The AssessorAgent generates questions from the same knowledge base it is testing. A sufficiently shallow knowledge base may produce questions that only require recall of stated facts. The four-level difficulty distribution and the analysis-weight multiplier (2.5×) mitigate but do not eliminate this risk.

Search engine fragility. The DuckDuckGo engine uses HTML scraping and is the most fragile engine in the aggregator. Rate limiting and HTML structure changes can silently reduce result quality. The quality scoring and multi-engine redundancy provide resilience, but engine-specific failures are not surfaced prominently to users.

Embedding model dependency. All semantic retrieval in the memory store and vector store depends on the consistency of the configured Ollama embedding model. Changing the embedding model between sessions renders prior embeddings incompatible. There is currently no re-embedding migration path.

Validation pipeline coverage. The eleven validation phases do not cover all correctness dimensions. A system can pass all phases and still exhibit incorrect business logic, security vulnerabilities, or performance pathologies. The phases verify structural and runtime correctness, not semantic correctness or security properties.

LLM-dependent sub-topic decomposition. The initial sub-topic decomposition is a single LLM call with no validation. If the LLM decomposes a topic in a way that misses critical sub-domains, subsequent learning cycles will be bounded by that initial decomposition. There is no mechanism to add new sub-topics after the first cycle.

13

Future Work

Reference confidence feedback loop. The SQLite reference confidence dataset is currently observational. Closing the loop—feeding high-confidence source domains and engines back into the aggregator's ranking logic as priors for related topics—is a natural next step. This would create a seventh feedback loop: sources that historically improve mastery would be preferentially retrieved on similar future topics.

Dynamic sub-topic graph. The current implementation fixes the sub-topic set at cycle one. Adding a mechanism to identify and add new sub-topics when gap analysis reveals structural gaps outside the existing decomposition would improve coverage for topics with non-obvious dependencies.

Knowledge ontology. The memory graph currently stores episode records (what happened, when, what was the outcome). Extending it to a full topic ontology—structured relationships between topics, sub-topics, and implementation patterns, not just similarity-based retrieval—would enable more targeted context injection and explicit prerequisite reasoning.

Domain-specific validation. The validation pipeline currently applies the same eleven phases to all implementation languages and frameworks. Specializing phases by domain—different contract checks for Go interfaces versus Python protocols, different startup verification for containerized services versus CLI tools—would reduce false failures and improve repair specificity.

SKILL.md interchange standard. Formalizing the SKILL.md format as a published interchange standard for AI agent knowledge transfer, with a schema, versioning, and a registry, would allow skill artifacts produced by this system to be consumed by any agent platform that adopts the standard.

External mastery validation. Integrating an external evaluation layer—a separate agent instance, a domain-specific benchmark, or a human-in-the-loop review protocol—as an optional post-processing step would address the self-referential validation limitation and provide externally verifiable mastery claims.

14

Conclusion

This paper has described a multi-agent system that autonomously acquires, measures, and applies domain knowledge. Its central contributions are:

  1. A closed-loop learning cycle with eight feedback mechanisms—including gap-directed search, mastery-gated focus, cross-session semantic memory, and structured lesson accumulation—that produces measurable convergent knowledge acquisition across cycles.
  2. A mastery scoring framework based on multi-level weighted self-assessment that distinguishes systems capable of recall from those capable of analysis, and applies exponential blending for score stability.
  3. A multi-modal research engine spanning ten concurrent search engines with Whisper video transcription, quality scoring, injection-filtered synthesis, and DOI enrichment.
  4. An eleven-phase implementation validation pipeline that produces independently verifiable, startable code artifacts through directed repair with external bug-triage integration and cross-session lesson memory.
  5. A fully local semantic memory architecture using Ollama embeddings that enables cross-session transfer learning without model fine-tuning or external API dependencies.
  6. A MITRE ATLAS-mapped prompt injection defense applied at data ingestion, before any LLM prompt construction.
  7. A skill export protocol producing structured SKILL.md artifacts that transfer completed learning sessions as loadable context packages for downstream AI agents.

The system demonstrates that meaningful, measurable knowledge acquisition—at the sub-topic level, across multiple learning modalities, with validated implementation output—is achievable through a well-designed feedback architecture over commodity language models and local inference infrastructure. No specialized training, fine-tuning, or cloud-dependent embeddings are required. The core loop is composable, extensible, and operates entirely within a self-contained deployment.

—

References

  1. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv:2305.16291.
  2. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. NeurIPS 2022.
  3. Novikov, A., et al. (2025). AlphaEvolve: A Learning Framework to Discover Novel Algorithms. Google DeepMind Technical Report.
  4. Schmidhuber, J. (2025). Self-Referential Basis of Undecidable Dynamics: Darwin, Gödel, Turing. arXiv:2504.17982.
  5. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020.
  6. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023.
  7. Significant Gravitas. (2023). AutoGPT: An Autonomous GPT-4 Experiment. GitHub repository. github.com/Significant-Gravitas/AutoGPT.
  8. MITRE Corporation. (2023). MITRE ATLAS: Adversarial Threat Landscape for AI Systems. atlas.mitre.org.
  9. Biryukov, A., & Dinu, D. (2016). Argon2: The Memory-Hard Function for Password Hashing and Other Applications. Password Hashing Competition, final document.
  10. M'Raihi, D., Bellare, M., Hoornaert, F., Naccache, D., & Ranen, O. (2005). TOTP: Time-Based One-Time Password Algorithm. RFC 6238, IETF.