Last updated: 2026-10-05

M
Masters level
APL
Applied / Methodological — Knowledge with a 5–10 year half-life — stable practice

From Semantic Search Towards Cognitive Agency: The Accidental Architecture of Memsearch

A personal knowledge tool that acquired several of the functions a cognitive architecture needs, and what that does and does not show

Memsearch did not begin as an attempt to build an agent. It began as a way of finding material across one person's working files and conversation histories, and it is still, at its core, a retrieval tool. As functions were added, though, its design started to resemble the cognitive cycle this module uses as its shared diagram. It senses its environment, keeps several kinds of memory, abstracts patterns across them, proposes connections, checks some of those connections against outside sources, revises them, and changes its own workspace.the diagram is a classic control loop

This page treats memsearch as a case study rather than a showcase. It is a work in progress. Its idea-synthesis pipeline is experimental, and the project's own documentation says that its output is speculative machine synthesis, to be investigated rather than cited. Nothing here claims that the system has been validated, and it is not yet a fully autonomous agent. The page asks where a search tool becomes a cognitive system, and where a cognitive system becomes an agent, and it uses the project's real behaviour to make both questions concrete.

What Memsearch Currently Does FoundationalKnowledge that endures for decades — core principles

Memsearch chunks and embeds the files of several tracked project directories, every conversation transcript from the author's coding assistant, including those from sub-agent runs, and a local library of PDFs. It stores the embeddings in a local vector database and builds a knowledge graph over them, including a map of nearest neighbours and two kinds of gap. Interpolation gaps are pairs of nodes that are related in embedding space but have nothing connecting them. Extrapolation gaps come from each cluster's outer edge: a small region spanned by several of its most outlying members reaches out in several directions, and the nearest existing node in each direction is a candidate for a bridge. A local web interface shows the graph and a board of candidate ideas and projects.

On top of the retrieval layer sits an idea-synthesis pipeline with seven commands. synthesize proposes cross-domain ideas from the graph's interpolation gaps, and synthesize-region explores the extrapolation regions in one of three modes: point, which treats each sampled direction's nearest neighbour as an ordinary connection; region, which describes the whole region without a concrete anchor; and combined, which grounds the region description in those nearest neighbours. Only point mode has been run end to end so far. verify takes a report's open questions to a headless instance of a hosted model with web search, and writes the findings back into the report. refine revises an idea against those findings. synergy looks for shared techniques or dependencies between live ideas. consolidate rolls related ideas into a named parent project, and decompose breaks a large idea into sub-projects with stated interfaces. Reports produced by synthesis are mined back into the store, so they become part of the material that later cycles search and draw on.

Three Tiers of Material Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The material the system searches is not uniform, and the tiers the site uses for knowledge decay[the half-life of knowledge] give a useful way to see the difference. Conversation transcripts are ephemeral. Each records a particular exchange, usually about a particular problem, and it is rarely consulted again once the problem is solved. The project artefacts the system indexes, such as specifications, papers, code, and notes, are closer to the foundational tier, because they keep their value over years. The method itself, meaning the way the pipeline is organised and the reasons for its design, sits at the applied level. It is used in practice and revised as it is used, but it is not yet settled.

The tiers matter for the system's behaviour. A search that returns a transcript from two years ago should be read differently from one that returns a specification, and the store does not currently mark the difference. The ephemeral-to-foundational pathway material describes how a quickly acquired idea can be promoted to a durable one, and memsearch does not yet do that promotion. It indexes both kinds of material with equal weight.weight is not relevance

A Nightly Cycle Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

A scheduled task runs memsearch once a day at 03:00. Its job re-mines the tracked project directories and the conversation transcripts, fetches new external sources for the PDF library, removes the chunks of any file that no longer exists on disk, and rebuilds the knowledge graph. Once the indexing and graph rebuild are finished, the same run goes on to synthesis. It proposes ideas from the graph's gaps, refines each one with the local model, mines the finished report back into the store, and scans the ideas for synergies. The overnight work is therefore both maintenance of the representation and generation, and the generated material is written back into the store that the next night's run will index.

The time was chosen for convenience, since a job that runs while the machine is idle does not disturb daytime use. The schedule nonetheless produces a resemblance to the sleep-based consolidation described on the sleep consolidation page. The system takes in experience during the day, in the form of files written and conversations held. It reorganises that material overnight, extracting the patterns that the embeddings and graph carry, and it removes material that is no longer there. The parallel with the biology is partial, and it should be read as a resemblance of timing and function rather than of mechanism. Nothing is replayed, and nothing is retained because it was reinforced. The pruning removes sources that have been deleted, not weakly connected material, so it is closer to clearing away the leftovers of a day than to the synaptic downscaling that Tononi and Cirelli propose. The parallel is this page's own reading, not something the sleep page claims about software.cf. biological consolidation via replay, not just timing

Mapping the System onto the Cognitive Cycle Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The table below is an interpretation. It reads memsearch's commands against the stages the cognitive-cycle page sets out[the cognitive cycle]. The project was not designed around that model, so the correspondences are the author's reading of the code, and the last column says how strong each one is.

Cognitive functionMemsearch mechanismInterpretation
Sensemine, mine-convos, and source ingestionAcquiring material from an environment made of project files, conversations, and documents
PerceptionChunking, parsing, and metadata extractionTurning source material into units the system can process
Short-term contextThe current query, the records it selects, and the candidate idea under examinationThe immediate working situation of one operation
Episodic memoryStored project files and conversation historiesRecords of particular past work and interactions, kept with their source
AbstractionEmbeddings, clusters, and nearest-neighbour relationsSimilarity structure drawn across many records
RepresentationThe vector store and the derived knowledge graphThe system's current organised account of the user's material
ThinkSemantic retrieval and synthesis over selected materialFinding relations relevant to a query
ImagineCandidate ideas proposed from gaps between clustersPossibilities not explicitly present in the source material
PredictNo dedicated stageConsequences of a candidate idea are not modelled explicitly, so this stage is largely absent
Verifyverify, which searches the web for the open questionsTesting an internally generated idea against sources outside the system
Reflect and reviserefineChanging a proposal in response to what verification found
Coordinatesynergy and the reconciliation of several decompositions in decomposeRelating independently generated analyses to one another
Actconsolidate, decompose, and the idea boardChanging the persistent workspace, not only returning text

The Predict row is the one most likely to be over-read. A reader might take the decomposition step, which states interfaces and dependencies, for a prediction of consequences. It is closer to a plan for a project than to a forecast of what the project will do. For the purposes of this module, the gap is that memsearch has no stage that runs a candidate forward and compares outcomes.plans are not predictions missing the simulation step

Not a Pipeline Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The commands are usually run in a sequence, which makes it tempting to describe the system as a pipeline: mine, search, synthesize, verify, refine, act. That sequence is a useful way to operate the tool. It is a poor account of what the system is doing, because the important feature is the feedback. Once a generated report or a project definition becomes part of the user's files, a later mining run ingests it, so action changes the environment that will be sensed next.

Environment (project files, conversations, PDFs)
    |
    v
Sense and ingest
    |
    v
Persistent memory and representation (vector store, graph)
    |
    v
Retrieve, associate, and synthesise
    |
    v
Generate candidate idea
    |
    +------> external verification
    |                |
    |                v
    +----------- revision and refinement
                     |
                     v
          project or knowledge action
                     |
                     v
          changed environment --> mine again

This is closer to the cycle than a conventional retrieval application is. The cognitive-cycle page makes the same point about biological and artificial agents: the stages are densely connected, action feeds back into the environment, and plans are revised by results. The loop in memsearch also carries a risk, discussed below, that the same feedback can amplify a weak idea.the classic control loop

Not One Memory but Several Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

A vector store is often described loosely as an agent's memory. Memsearch shows why that description hides more than it reveals. It holds several things that perform different memory functions.

  • Source memory. The original files and transcripts remain the primary record. They preserve particular events and discussions, which is closer to episodic memory than to abstraction.
  • Semantic or abstract memory. Embeddings and clusters discard much of a source's exact form and keep patterns of similarity. They work more like abstraction than recollection.
  • Relational memory. The knowledge graph records connections between items that may come from different projects and conversations.
  • Working context. A query, a candidate idea, a verification report, or a decomposition attempt forms a temporary working set for one operation.
  • Prospective memory. The idea board records matters intended for later attention. It stores intentions and unfinished possibilities, which are not memories of the past.

The cognitive-cycle page separates episodic experiences from generalised abstractions and from the immediate working situation. Memsearch makes the same separation in its storage, and the distinction matters because each kind is revised in a different way. Source material should not change because a report about it was written. Generalisations change when the material changes. Prospective entries change when a person decides they are finished.raw logs vs. learned rules vs. current plan

Gap-Finding as Constructive Imagination Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The most interesting feature is the attempt to generate ideas from gaps between clusters. The project's own description of an interpolation gap is precise: a pair of nodes that are related in embedding space but have nothing bridging them. Extrapolation gaps are a different kind of candidate. They lie past a cluster's edge, along directions the cluster's outermost members point toward, so they are regions to explore rather than pairs to connect. That is a geometric fact about the representation, and it says nothing about whether the two areas should be connected.

An embedding-space gap is therefore not an undiscovered fact, a valid research gap, evidence that two concepts belong together, or proof of novelty. The system does not discover ideas latent in the vector space in the way a search engine finds matching records. It constructs candidate relations from the geometry of its representation. Those candidates may be productive, trivial, mistaken, or artefacts of the embedding model. Until they are investigated separately, their status is imaginative.

This fits the cognitive-cycle account of imagination, which treats it as recombining fragments of what is already known into situations that have not occurred, rather than extrapolating the present forward. It also fits the project's own warning that generated connections are leads for human attention, not conclusions.recombine, don't just predict

Verification Outside the Generating Loop Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The robust-agent page argues that verification must come from a path that does not run through the component being verified[designing auditable agentic systems]. Memsearch applies this in a concrete form. The local model that writes synthesis reports has no web access, and it is instructed to name anything it cannot check as an open question rather than assert it either way. A separate step takes those questions to a hosted model with web search, and writes the findings back into the report. The generating stage and the checking stage are separate processes with different capabilities, and that separation is the design point.

Separation of this kind does not make verification independent. Two models can share training sources, assumptions, and failure patterns, and a checking stage can accept the framing that the first stage supplied. Verification in this system should therefore be split into distinct checks, each with its own question.

CheckQuestion
Source verificationDo supporting sources exist?
Claim verificationDo those sources actually support the claim?
Novelty verificationIs the alleged connection already well established?
Contradiction searchWhat evidence counts against the proposal?
Provenance checkWhich parts came from source material, and which were generated?
Human judgementIs the idea worth pursuing in this context?

The system performs the first and some of the second, and the last belongs to the person. The middle checks are where the design is weakest, because a web search can find a source that loosely resembles a claim without supporting it.

From Retrieval to Agency Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

A program does not become an agent because it calls a language model. The stronger agentic characteristics appear when a program maintains persistent state, works toward an object or goal, chooses among possible actions, observes the effects of those actions, updates its model, and continues the cycle. Memsearch implements several of these, but the high-level sequencing is still started by a person, one command at a time.

LevelCapabilityMemsearch example
Retrieval systemFinds relevant stored materialSemantic search
Representational systemOrganises relations among stored materialEmbedding space, graph, and clusters
Cognitive support systemGenerates and evaluates possible connectionsSynthesis, verification, and refinement
Agentic workflowSelects and performs operations that change persistent stateConsolidation, decomposition, and project organisation
Autonomous agentChooses its own goals and continues acting without promptingNot present; not needed for the tool to be useful

Memsearch sits in the agentic-workflow row. It is a cognitive tool with agentic workflows rather than an autonomous agent. That intermediate category is both accurate and useful for teaching, because it shows which functions can be built incrementally, and which would need a different kind of control to be trusted.

Human and System Roles Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The tool is not simply automating the production of ideas. Its division of labour looks like this.

System contributionHuman contribution
Searches across more material than anyone can hold in working memoryRecognises which findings are significant
Detects weak or distant similarityJudges whether the similarity is meaningful
Generates possible bridges between areasSupplies disciplinary and personal context
Repeats verification and revisionDecides when the evidence is adequate
Tracks candidate ideas and their relationsChooses goals and priorities
Decomposes work into possible projectsAccepts responsibility for undertaking them

The machine supplies persistence, breadth, association, and repeated critique. The person supplies motive, situated judgement, and responsibility. This makes memsearch a concrete case of the many-eyes argument in the generative AI curriculum page: the value lies in several different perspectives meeting, and the responsibility for the outcome stays with a person.

Failure Modes and Epistemic Risks Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

A case study is only as useful as the risks it names. Several are already visible in the design.

  • Representation failure. Embeddings can make superficially similar passages look related while separating concepts that are connected but described in different vocabulary.
  • Memory pollution. Synthesis reports are mined back into the store. A later search can return a generated report beside original evidence, with no structural difference between them. Over repeated cycles, speculation could become hard to tell apart from source material.
  • Self-reinforcing synthesis. If a generated idea is re-ingested, a later synthesis may treat the earlier speculation as evidence that a relation is important. The loop that makes the system cumulative is the same loop that could amplify a weak idea.
  • Verification theatre. A verification stage can look rigorous while checking only details that are easy to search for, or while accepting sources that loosely resemble a claim.
  • Goal drift. A system may optimise for ideas that look interesting rather than ideas that are useful, feasible, or true. Nothing in the pipeline currently measures usefulness.
  • Salience bias. Heavily represented projects can dominate the embedding space, while small, recent, unusual, or poorly documented projects are hard to retrieve.
  • Privacy and boundary failure. The tool is local-first, but verification sends queries derived from private material to an external service, so some exposure remains.
  • Premature project formation. Consolidation and decomposition can give a speculative idea the look of maturity by wrapping it in a project name, work packages, and interfaces.

These risks do not invalidate the project. They define the controls and evaluation questions that the next version needs. The brittleness described on the agents-that-fail page applies here too. The store holds only what has been mined, so a question far from that material may be answered with confidence and little basis, which is an inference from the architecture rather than a measured result. The same applies to output that misrepresents its own status when a generated report sits beside original evidence with nothing marking the difference.

What Would Need to Be Evaluated Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

At present the project has no evaluation against a standard that would make its output a sound basis for research conclusions. A useful evaluation would ask whether the candidate ideas that survive verification are judged useful by people who know the domain, whether verification finds genuine errors in generated claims, whether generated material can be told apart from source material by someone reading the store, and whether the system's suggestions change what its user actually does. Usage counts alone cannot answer those questions.

Questions the Project Leaves Open Applied / MethodologicalKnowledge with a 5–10 year half-life — stable practice

The project is unfinished, and its open questions are a productive part of the case.

  • Should generated material be stored apart from primary sources?
  • Should every derived idea carry a complete record of its provenance?
  • How should confidence change after verification, and in which direction?
  • Can the system be made to seek disconfirming evidence actively, rather than only checking claims it has already made?
  • When should refinement stop?
  • Who or what decides which idea receives attention?
  • Can the system recognise duplication with an established field described in unfamiliar language?
  • Should actions that change persistent state require explicit human approval?
  • How can usefulness be evaluated without rewarding novelty alone?
  • Can the system explain why two clusters were treated as distant or as bridgeable?
  • What should happen when local knowledge and external evidence conflict?
  • How should forgetting, expiry, and correction work for generated material?

These questions turn an incomplete implementation into a design enquiry, and they match the questions the auditable-agents page raises about where checkpoints belong.

Closing FoundationalKnowledge that endures for decades — core principles

Memsearch did not start out as an attempt to implement a cognitive architecture. It started with the practical problem of finding and connecting material scattered across one person's work. The need to take in experience, keep several representations of it, retrieve relevant episodes, abstract patterns, imagine connections, test them, revise them, and turn some of them into future action has nonetheless produced many of the same components. That convergence does not show that the system thinks, and it does not show that every search tool with a language model is an agent. It does suggest that cognitive architectures sometimes emerge from confronting the functional problems that cognition solves, rather than from imitating cognition directly.