October 2, 2026
The Model Is No Longer the Whole AI System.
What this week’s Deep Learning Weekly tells us about memory, agents, architecture, and evaluation

By José. T Guevara
3 min read
By José Tomás Guevara Calderón, Independent Researcher
This weekend's edition of Deep Learning Weekly, curated by Miko Planas, deserves attention for something larger than any individual model announcement.
Credit belongs to Planas and Deep Learning Weekly for assembling the developments discussed here. My purpose is not to reproduce that publication, but to examine what I believe connects several of the developments it highlights.
And that connection is increasingly difficult to ignore:
The frontier problem in AI is becoming an architecture problem.
Bigger Models Are Not Enough
For several years, much of the AI conversation revolved around model capability: more parameters, longer context windows, better benchmarks, faster inference.
Those things remain important.
But an autonomous agent operating for minutes is fundamentally different from an intelligent system expected to function coherently across hours, days, interruptions, changing objectives, tool failures and enormous histories of previous interactions.
The question changes from:
How intelligent is the model?
to:
How well organized is the system around the model?
That distinction is fundamental.
The Memory Problem Makes This Visible
One particularly interesting development highlighted in the newsletter is research into Just-in-Time Memory for LLM agents.
The idea described in the material is important: instead of deciding permanently what an agent should remember when an experience occurs, retain richer historical information and curate the relevant memory when a future task establishes what is actually needed.
That represents an important conceptual shift.
Memory is no longer merely storage.
It becomes an architectural process involving selection, relevance and context.
That resonates strongly with the problem we have been examining in Salomon-Prime (SP).
SP approaches persistent AI from a broader architectural perspective: an intelligent system needs mechanisms for organizing persistent state, active context, memory, objectives and operational continuity rather than repeatedly reconstructing itself from increasingly large conversational histories.
Importantly, this remains an experimental hypothesis, not a demonstrated superiority claim.
It should be tested.
Token Efficiency Is Also Architectural Evidence
The same issue highlights work on reducing token consumption during long-running agent operation.
This matters economically, but I believe there is a deeper implication.
If an intelligent system must continuously reload enormous amounts of previous information simply to maintain functional continuity, token consumption may reveal something about the organization of the architecture itself.
An efficiently organized system should ideally determine:
What must remain persistent? What can be forgotten? What should be retrieved? What constitutes current state? What information matters to the present objective?
These are architectural questions.
And Then Comes Evaluation
This is where FCSEP, the Functional Cognitive Systems Evaluation Protocol, becomes relevant.
Building persistent agents is only half the problem.
We also need to know what happens when they are disturbed.
A capable agent under ideal conditions tells us relatively little about its resilience.
An evaluation should expose the system to controlled perturbations:
baseline → perturbation → degradation → recovery → continued operation
Can it preserve functional state?
Can it recover after interruption?
Does memory remain reliable?
Do errors accumulate?
Does token consumption explode as history increases?
Does the system continue pursuing the correct objective?
FCSEP is being developed precisely to frame questions of this kind systematically.
The Experiment I Would Like to See
Rather than arguing abstractly that one architecture is better than another, we should perform a controlled comparison.
Give two agents:
- the same underlying model,
- the same tools,
- the same information,
- the same tasks,
- and the same computational environment.
Let one use a conventional agent architecture.
Let the other operate through an SP-style persistent architecture.
Then measure:
task success, memory precision, state continuity, token consumption, latency, recovery, error propagation, tool efficiency and total computational cost.
Finally, perturb both systems and evaluate their behavior using FCSEP.
If SP provides no meaningful advantage, we learn something.
If it does, we learn something considerably more interesting.
This Is Bigger Than Salomon-Prime
The real lesson I take from this week's Deep Learning Weekly is not that external research has somehow "validated" SP.
That would be an unjustified claim.
The interesting development is that independent researchers and engineering teams are increasingly addressing memory organization, context management, agent controls, evaluation, efficiency and long-horizon operation.
Those problems point toward a broader transition:
AI research is gradually moving from models toward systems.
The next leap may therefore come not simply from another larger model.
It may come from learning how to organize models into systems capable of maintaining memory, continuity, resilience, purpose and coherent operation over time.
That is the territory Salomon-Prime intends to explore.
And FCSEP gives us a way to ask whether it actually works.
Acknowledgment: This commentary was inspired by Deep Learning Weekly and its coverage of current AI industry and research developments. Credit for the newsletter's selection and presentation belongs to Miko Planas and Deep Learning Weekly. The interpretations concerning Salomon-Prime and FCSEP are my own.
Why I framed it this way
I deliberately did not claim that the newsletter validates SP. Instead, it establishes independent convergence around problems SP addresses.
For more information contact me: josetguevara108@outlook.es