The autonomous agentic substrate for the household, grown on Buzz.
Foundation-AI is being built as a governed operating system for agentic AI: a platform for fleets of persistent, autonomous agents operating as a coordinated Hive to execute real-world work across enterprise and consumer environments. It enables scalable autonomous operations with integrated capabilities for intelligence, memory, governance, security, knowledge management, workflow execution, and value exchange.
Foundation-LifeStyle, part of the Foundation-AI series, is built on Buzz, Block’s open-source agent-native network, extended from the outside into an autonomous agentic substrate for the home. Buzz gives the twins a shared, open place to live and act as members of the household rather than bolted-on bots, and Foundation-LifeStyle adds what a home needs on top without changing Buzz to do it. What sits on top is not a thin wrapper. It is a substantial system in its own right, a full autonomous substrate, while the Buzz underneath stays exactly as Block ships it.
A home runs on a stream of small work that nobody has time to watch for, and the people sharing that home keep things from each other for perfectly good reasons. Both of those are why one AI for the whole family is the wrong shape.
It runs today in a real home. Each person has their own AI twin, called Honey and Pollen in the house we follow here; the household has one of its own, named Buzz after the network it lives on; and a fourth agent, Bee, looks after the software so that nobody in the household has to.
A house generates a steady stream of small work that nobody has time to watch for. A warranty that runs out in March. A filter that should have been changed. A recall notice that was mailed to whoever lived here before you. An air conditioner that is taking nine minutes longer to cool a bedroom than it did last summer, under weather you could fairly compare.
None of it is difficult. All of it needs somebody paying attention at the right moment, months after anyone last thought about it. That is the job this system does, and it is a different job from answering questions faster.
The reason it cannot be one AI serving the whole family has nothing to do with how capable the model is. A household shares a building, a calendar and a bank account, and deliberately does not share everything else. Presents. A diagnosis. What somebody earns. Which of you is having a hard month. A single AI holding one pot of memory for all of that has to be trusted to answer carefully, and being trusted to is a promise rather than a property.
To be clear about what that does and does not mean: most of what a household knows is shared, and this treats it as shared. The calendar, the house, the shopping, the trip, what is for dinner and who is picking up whom. The separation is for the short list, not the long one. A family that had to authorize every ordinary fact about itself would be unbearable to live in, and that is not what any of this is for.
So this is built the other way round. Each person gets their own twin. The house gets one of its own. A fourth looks after the software and holds nothing personal at all.
A twin is not something watching you. It stands in for you, which means it knows the things you would otherwise be carrying in your head, so that when it deals with the world on your behalf it answers the way you would have. The house twin does the same for the building: what every machine is, how hard it has been working, when its warranty ends, and which papers and photographs in this house are about it.
What each twin knows lives in a separate file with a different owner, and the operating system refuses the others at the door. That is not a policy anybody has to remember. It is the same mechanism that stops one login reading another login’s documents on a shared laptop, used deliberately in a place where the industry has mostly been relying on good intentions.
The work itself is done by standing jobs called Scouts, which run for weeks or months and interrupt you once, at the moment it is genuinely your call. This paper walks through nine things they look after in a real house, and a tenth that the household adds itself.
One assistant can serve one person. A home needs an autonomous substrate.
None of this runs on a private copy of anything. It runs on Buzz, an open network that Block gives away, used exactly as it ships.
Buzz is an agent-native network: a shared place where people and their AI agents are members side by side, each with its own name and signature, rather than software posting under a human’s account. Block built it, gives it away, and anyone can run it and read every line of it.
Foundation-LifeStyle does not fork it, patch it, or slip a changed copy inside a product. It takes the network as it ships and adds what a household needs from the outside, without touching what is underneath. That restraint is deliberate, and it buys three things a family should care about.
Nothing is locked in. The foundation stays the real, open one, so a household is never stranded on a private version that only one company can keep alive. What Block improves upstream, the house gets.
Nothing is hidden. Because the ground is unchanged and open, you can see exactly where the neutral network ends and the product begins, and the parts that hold your family’s information are on the side you are free to inspect.
Nothing is bolted on. The agents are first-class members of the network from the first moment, not add-ons wedged in under one person’s account, which is what lets each one carry its own identity, keep its own separate memory, and act with authority that can actually be checked.
Everything else in this paper is that outside layer, and it is not a small one. The twins, their separate memories, the standing jobs, the settlements between them: it is a major autonomous substrate, grown whole on top of a Buzz that is never touched, without forking a single line of it.
A handful of named parts do the actual work behind every capability here. The same few keep turning up, so they are worth knowing by name.
Scouts are standing jobs. You hand a Scout something to watch for, a warranty that lapses in March, a ticket under a price, a filter due for a change, and it watches on its own schedule, for months, acting only when there is a reason. It survives restarts and outages because it writes down where it got to after every step. Most of the useful work in a house is a Scout.
Bee is the one agent that looks after the software itself, so nobody in the household has to. It notices when something has broken, prepares the fix, and hands anything that touches money, identity, or privacy to a person to approve rather than shipping it alone. It never grades its own work: an independent check has to reproduce a fix before it counts.
PACT is how one twin asks another to do something, Honey asking Pollen, say. It is a private contract between them, so one twin can get a piece of work done for another person in the house, or settle up for it, without ever seeing that person’s private life. A request is only ever a request, never a standing permission.
HIVE is how the twins share what they have learned, but only once it is proven. A guess never travels; only a verified outcome becomes a signal another twin will act on, and it expires when it goes stale. This is what lets several twins agree on something without pooling their private lives.
LogOS is the window the household’s administrator has into all of it. It shows that things are running and being kept in repair, what worked and what needs a person, while showing none of anyone’s private content. It is oversight without surveillance.
One idea sits under all of them, and it is the real core of the substrate. Everything these parts do is written down as connected facts, each action linked to what caused it and what proved it, and that growing web of linked, permanent records is the thing everything else stands on. A Scout’s watch, a fix from Bee, a contract carried over PACT, a proven result shared through HIVE, an entry LogOS can show the household: each is a point on the same graph, and nothing happens without leaving one. It is what lets the system always say how it reached a decision, and why no part can quietly rewrite what already happened. Engineers call it the graph; for a family it just means nothing is done without a record of why.
Every capability in the rest of this paper is some arrangement of those parts, standing on that graph, on the open ground of Buzz.
This does not begin with AI. It begins with the ordinary pile of things in a house that needed doing and did not get done in time.
The washing machine that failed three weeks after the warranty ran out. The insurance that renewed at nearly twice what it cost last year. The subscription nobody remembers signing up for. The part that was cheap in April and unobtainable by September. Nobody thinks of those as AI problems, and nobody is to blame for any of them.
What they have in common is that each one was findable in advance by somebody who was watching, and nobody was watching, because watching is dull and continuous and people are bad at it. An assistant does not fix this, because you have to remember to ask, and remembering to ask is the part that failed.
So the unit of work here is not a question and an answer. It is a standing job, called a Scout, and it is the single thing that makes this different from every assistant on the market. A Scout has a subject, a set of rules, a budget, a spending limit if it needs one, and a stated idea of what finished looks like. Some close in an afternoon. Others run for months, waking whenever something changes.
Underneath, that standing job is an agentic mesh: agents shaking hands over PACT, talking on a shared LogOS bus, and every move written down on the graph. Everything in the nine sections below is a Scout doing something. The areas are what a household happened to need. The Scout is the product.
One rule sits in front of all of it: a Scout is only created when a single answer will not do. If the question can be settled now, it is settled now and nothing is created. A system that turns idle questions into permanent chores is not being helpful. It is accumulating clutter that remembers.
The nine areas below are not a feature list somebody drew up. They are what one household actually needed, in the order it needed them, and each one is running in code today at a stated level of proof. The tenth is the point of the whole exercise: a household can add an area nobody anticipated without knowing anything about computers.
What makes them able to coexist in one house, rather than as nine separate apps each with its own account and its own copy of your life, is the substrate underneath. That is worth establishing first, because every section after it leans on the same handful of rules.
Six rules do the structural work. They are stated here once, in plain terms, because the rest of this paper is nine applications of them.
There are three database files, one for the household and one for each person, each owned by a different account on the machine. The obvious alternative is to put everything in one file, tag each row with whose it is, and filter on the way out. That works exactly as long as every future query remembers the filter, including code written next year by somebody in a hurry at eleven at night.
A separate file with a different owner is a boundary the kernel holds for you, and it keeps holding when the query is written by a tired person, by a different program, or at a shell prompt with the file open. The file mode is the enforcement. Read-only connection flags exist in the code as well, but they are checks a program applies to itself, and a program that applies a check to itself can stop applying it. The one guard a process cannot talk its way past is the operating system refusing the write.
The watching itself is continuous and the thinking is not. A standing job wakes when there is a reason, does one piece of work, writes down where it got to, and stands by, which is why a watch set in one month is still running months later across every restart in between.
The system does not trust its own configuration on this point. A startup check re-launches itself as each twin’s user account and genuinely attempts to write to a store that account should not be able to write to, and requires the attempt to fail. A successful write is the loudest finding that check can produce. The check also refuses to be graded on a curve. It has three outcomes rather than two - it passed, it failed, or it could not be determined - and in production the third is treated exactly as seriously as the second, on the reasoning that not being able to establish who owns the household store is not a milder problem than the wrong person owning it. Which machine it is running on is told to the check rather than guessed by it, so a development machine cannot quietly present itself as the deployed one.
Anything that runs while nobody is watching has to answer a harder question than "did it work". It has to answer "how would we know". So the component that records a verification refuses a verifier that shares an identity with the thing that acted. That check is made twice, deliberately. Testing only what the caller passes in leaves the obvious hole - a caller can pass one identity and record another - so the same independence test is applied again to the identity that actually got written down, and the two have to agree. Restarting a service is not evidence the service works; the repair is confirmed by re-running the exact probe that found the fault, which is a different question asked of the world rather than a report from the actor.
There is no branch anywhere in the repair loop that turns "I could not tell" into "resolved". Inconclusive escalates.
This rule shows up more often than any other below, and it separates a system you can leave running from one you cannot. When a sensor does not answer, the store records that it did not answer. It does not record a zero, and it does not carry yesterday’s reading forward as though it were today’s. Recording that is deliberately made no harder than recording a value, because a rule that costs more to obey than to break does not survive contact with a deadline.
Before anything is changed, the exact set of things it will touch is worked out and counted. The failure this prevents is small and ordinary: somebody says turn the heating down, three radiators are in scope, the person picturing it pictures one, and nobody counted out loud. An approval is an approval of a specific blast radius, so if any of those targets has moved since the plan was computed, the plan is refused rather than applied on a best-effort basis.
The approver is resolved against the register of who is a person, and an agent is refused. That is checked in three separate places, because a violation here is indistinguishable from correct behavior right up until the moment it matters.
A fact cannot enter the store without a row saying what produced it. Third-party content cannot be marked trusted by the code that ingests it, a retracted source cannot support a fresh claim, and when two sources disagree the disagreement is recorded as a fact in its own right rather than averaged away.
Search results are pointers, never answers. A hit carries an identifier and a score and deliberately carries no text at all, so composing a reply out of one is structurally impossible. To say anything, the system goes back through the identifier, re-checks both time axes, confirms the fact was not retracted, and loads its evidence.
No component can widen its own permissions. The repair loop cannot change the list of things it is allowed to repair, and changing that list is itself the highest escalation level, which is a human decision rather than a prompt the model can argue around. Good behavior is evidence for a decision a person makes. It is never the decision.
A voice assistant can already turn on an air conditioner. That is not the gap.
The gap is everything else about that unit. How many hours its compressor has run. Whether its warranty is still live. What its error code means, read against the photograph of its rating label that somebody took years ago and forgot. Which manual in the house belongs to it, which invoice, which service visit. A smart plug gives you a switch. It does not give the machine a history.
So each unit stops being an entity in a list and becomes an object with a past: what it is, where it lives, how hard it has been working, and which papers and pictures in this household are about it. The links are never invented. A manual is attached to a unit only when the document actually names it, and otherwise the record says the manual is not in the index, which is a different fact from the unit having no manual.
Reading is free. Knowing the living room is 25 degrees is not an act, and it needs no ceremony. Changing something is a different matter, and the boundary is drawn by absence rather than by permission.
No lock, alarm, garage or door entity is reachable through the house connection at all. Not gated behind a confirmation. Absent. The reasoning in the code is worth repeating, because it applies to far more than door locks: a confirmation is a check that a sufficiently confused caller can still talk its way through, and the boundary you can prove is the one made of what is not present.
For the things it can reach, the twin produces a proposal and stops. The module that works out what a machine should do imports no client that could call a service and holds no credential of its own. It makes one network read, produces an ordered plan with the exact arguments, and a person makes the change.
When a change does execute, the sequence is written to survive a crash honestly: revalidate the current state, durably record the proposal and the approval and an action marked in-progress, make the one exact call, then independently read the device back and record what was actually observed. Only a matching read-back is recorded as proven. A crash midway leaves an explicit in-progress record rather than an invented success.
An air conditioner does not fail. It gets slightly worse over a year, throws no error code, and crosses no threshold. Catching that means holding this machine’s behavior today against the same machine’s behavior months ago, under weather it can be fairly compared to.
The measure is effort against work done: what the compressor drew, divided by the temperature difference it actually sustained against the outside air. That ratio is flat for a healthy machine and rises as a filter loads or a coil fouls. The hard part is not the arithmetic, it is the comparison, because a unit works harder on a hotter day and that is not degradation.
Dividing by outdoor temperature would trade one distortion for another, so the code refuses to fit a physical model it cannot validate against six units in one house. Samples are bucketed by outdoor temperature in two-degree steps and by whether the machine was heating or cooling, and only matching buckets are ever compared. A summer three degrees hotter than the last one moves samples into different buckets rather than into a verdict.
On one comparison the unit had taken fourteen minutes to bring a bedroom to temperature the previous July, and twenty-three minutes a year later under conditions matched on outdoor temperature, humidity and the sun load on the west wall. Twice before that, the watch declined to say anything, because it did not yet have enough genuinely comparable days. Saying nothing was the correct output and was recorded as an outcome rather than as an absence.
What reached a person was a booking already filled in: the model number read off the rating-label photograph, the manufacturer service bulletin covering that exact pattern, the warranty end date taken from the purchase receipt in the document store, and the authorized dealer dropped into the one free slot on the shared calendar. One tap sent it.
Two corrections say more about how this behaves than any success does.
A handover note recorded that three of the six air conditioners had no local interface and could be reached only through the manufacturer’s cloud, and three machines had been written off on that basis. The claim had never been tested, and it was false: all three answer on the local network using a newer protocol, and only the old interface is missing, which is what made them look dead to anything that checked the old way. The new reader was then cross-validated against the three units the existing integration could already read, at the same instant, through two independent code paths.
The second is smaller and more instructive. A module that proposes labeling unnamed units once asserted that the units resting on the weakest evidence were exactly the ones it could not reach. That is a tidy sentence, and it was false by the module’s own confidence table. It now computes the comparison and names the counterexample out loud rather than listing three steps and hoping nobody checks.
The most revealing information in a house is not in its documents. It is the minute-by-minute record a body produces.
The intended capability here is a long-term baseline per person, built from signals the house already emits, with change detection against that person’s own history rather than against a population average. It is the feature with the most obvious value and the most obvious way to go wrong, because it is health inference from ambient data about people who cannot meaningfully consent to every future inference.
So the disclosure rule was built before the detection was. Findings surface to the person before anyone else, without exception. In the code that is not a preference expressed in a settings page; it is the order the components were written in, and the detector sits behind the gate rather than beside it.
What the detector can do is narrow on purpose. It compares samples that are already in that person’s own private store, and it returns one of three things: not enough history, stable, or a hypothesis about a sustained change in one named signal. That is the whole vocabulary.
What it cannot do is longer, and each item is a refusal in code rather than a guideline. It cannot collect a signal. It cannot open the household store. It cannot compare two people. It cannot use a population norm. It cannot label illness, decline, cognition or wellbeing. The list of signals it may ever consider is closed, and a request naming a signal outside that list is refused, so a future component cannot quietly widen what is gathered.
One of the four intended signals is not collectable at all today, and the reason is mundane: the phones use private wireless addresses, so there is no per-person association to read. The system reports that as an absent signal rather than substituting something that correlates.
Health is where the separate files stop being an architecture diagram and start being the reason any of this is safe. Your medication, your allergies and whatever your watch made of last night sit on your side, in a file the other twins cannot open. What crosses to the house is a constraint somebody set, and never the reason behind it.
Mail is handled the same way and is the most privacy-sensitive capability in the system, because a mailbox is one person’s correspondence with people who never agreed to be part of any of this. Credentials reach only that person’s own twin. Nothing about mail is ever written to the household store: not a summary, not a sender, not a count. Messages are opened read-only, because merely fetching one would otherwise mark it as read and quietly destroy the unread state that person uses to decide what still needs them.
A household folder is the least glamorous data in the house and the most useful, and it is also the place where an AI is most likely to sound right and be wrong.
The whole design here turns on one requirement: when the twin tells you what your warranty says, it points at the page and the sentence rather than paraphrasing something it half remembers. Everything else follows from making that checkable.
Text is cut into chunks that never cross a page, slide or paragraph boundary, so a chunk’s page is the page rather than mostly the page. The usual approach packs text to a target size and records where it started, which produces citations that are right most of the time. A citation that is right most of the time is worse than none, because the reader cannot tell which kind they are holding.
Where a format has no real pages, it says so. A word processor document is paginated by whatever renders it, from live font metrics, and two applications disagree about where the breaks fall. Emitting page 4 for one of those would be fabricated provenance, which is a number that looks checkable and is not, so those documents cite their paragraph index and character offsets instead. Every offset is verified to actually index the extracted text when the document is built, rather than trusting seven extractors to have got their arithmetic right.
The table of formats includes the refusals as entries rather than as gaps. A photograph is a registered extractor that always declines, which is the point: a scan named Furnace-Warranty.jpg produces a record marked unreadable with a stated reason, so the twin can say there is a scan called Furnace-Warranty.jpg that it cannot read. That is a useful true statement, and it is much better than a household discovering by accident that its photographs were silently ignored.
Optical character recognition is deliberately the last resort. A document carrying real text should be read as text, because extraction is exact and transcription is a good guess with a confidence score. So it runs only where the primary extractor already failed, which means every transcribed document is one that would otherwise have been recorded as unreadable, and turning it on can never lose information.
In this household that was 130 documents out of 402: seventy-five PDFs whose fonts carry no usable encoding, twenty-eight with no text layer, eighteen that are simply photographs. Tax packets, bank passbooks, utility bills and receipts, which is to say the documents somebody is most likely to ask about, and the ones a search could not find because nothing but the filename was indexed.
It runs on the machine rather than in a cloud service, for a reason stated plainly in the code: a page image of a bank passbook is far more revealing than a mathematical summary of it, and sending one away would breach the spirit of the household’s own no-remote setting while satisfying its letter. Every transcribed passage is marked as transcribed, with its confidence, and that travels into the citation. The difference between what a bill says and what a transcription of a bill reads is small until the subject is money.
Removing a document from the twin’s memory is a one-way latch. There is no reverse gear, no second attempt, and no rebuild that brings it back. Which makes one line the most dangerous in the entire pipeline: the file is not in the mirror, therefore remove it. One failed sync, one unmounted drive, one wrong folder setting, and a household’s entire document memory is irreversibly gone from retrieval, unattended, at four in the morning.
A file that was not seen is not a file that was deleted. Four guards hold that line. If nothing at all was seen while the store holds live documents, the run fails and records nothing, deliberately, because letting a broken sync write absence records would let three broken nights in a row satisfy the fourth guard and delete everything. If more than a small fraction of documents are missing, bulk removal needs explicit human confirmation. If the same content appears at a new path in the same scan and nothing else matches it, that is a rename and is handled as one. Otherwise a document must be absent several times over a minimum span before anything happens, and until then it is simply not seen, which is unknown, and the twin keeps answering from it.
The scheduled run has four exit codes rather than success and failure, because an operator would act differently on each. One of them means the run completed with individual documents unread, which is the normal state of a real household folder and is explicitly not an alert.
The most seductive feature in this whole system is the one that turns a family’s photographs into a story, and it is the one that has been held back hardest.
The intent is a layer of named, bounded episodes over years of material: the week somebody’s mother visited, the summer the gym air conditioner kept failing, each with participants, places and evidence, so that questions can be asked the way a person would ask them. The deployed version derives 208 bounded episodes across 26 source albums, with no overlap inside an album and 6,198 explicit photo references.
Cutting a photo library into time-bounded groups is arithmetic. Naming one is not. The week somebody’s mother visited is a claim about a named person, her relationship to somebody in this household, her presence in a place, and the purpose of a gathering. Those are four separate assertions, and a system that produces that sentence out of a folder name has invented all four.
So the naming rule is mechanical rather than tasteful: every word in an episode name has to come from something the system can point at. That is what the index is for. Faces, objects and relationships are indexed, so a person, a place and an occasion become recorded facts with evidence behind them instead of inferences from a filename. Vision labels what each photograph was seen to contain, as an observation about that photograph. The relationships are the household’s own, because only the household can say who these people are to each other.
That is what makes show me the photos from her last birthday a question with an answer. The person resolves to somebody the household identified, the date resolves to photographs that carry it, and the occasion resolves to an album somebody in this house made and named. Every part of the answer points at something.
What the rule forbids is the sentence with nothing behind it, and the most common version is a summary quietly promoted into a claim. Food, tableware drawn from four of four hundred photographs is a statement about four photographs. It is reported with the number that actually contributed, rather than as a description of the day.
The supervision signal is the household. An album is a human decision: somebody in this house decided that these photographs belong together and gave the group a name. Nobody had to be asked for that labeling, and it is what bounds an episode. Episodes are therefore cut inside an album by time gaps and never across two albums, because merging two albums into one story is exactly where invention begins. Albums that overlap in time are reported as a fact rather than fused into a narrative.
Coverage is treated as a first-class fact rather than a footnote. The analysis was partway through, and the part that had been analyzed was not a random fifth of the library, it was three albums out of twenty-six. A layer that quietly described only what it could see would have reported a household that took no photographs before a certain year, and a hole shaped like an absence of events is indistinguishable from a hole shaped like an absence of data unless something says which it is. So every episode carries its own coverage statement, and a source can come back in five different states rather than two. Read, empty, absent, unreadable and malformed are five different answers, and an empty calendar is a genuinely different fact from a calendar that could not be reached.
Search over it is typo-tolerant, and it matches the album titles the household wrote, the people it identified, the dates the photographs carry and the labels on each one. Every read is revalidated against the exact current state of the photo library, and a stale result or no match returns unknown, never proof that an event did not happen.
The reason the rule is mechanical is written into the module. This project has already had to retract one fabricated memory note about this household. A confident episode name would be the same failure with a much larger surface, because it would be indexed, retrievable, and repeated back to the family as its own history.
Most of what a household knows was never typed into anything. It is in photographs, in video, and in a drawer of paper.
The scope is narrow on purpose: not the whole library, only albums somebody deliberately put on an approved list. Everything inside that scope is indexed and its metadata extracted, so what is in a photograph, who is in it, and when and where it was taken all become facts you can search rather than a pile of files with dates on them.
Video is the part people underestimate, and this is one of the most useful things the whole system does. Home video is not really a visual archive. It is mostly people talking, and the talking is the part worth searching. Transcribing it turns what did the builder actually say about the timeline into a question with an answer, and it does the same for every conversation anybody ever filmed in this house. Years of family video stop being an archive nobody opens and become searchable by what was said out loud in them.
Speech recognition on real home video is genuinely hard: several people at once, a room with echo, a television on, somebody half off camera. So the transcript keeps its confidence line by line, and the uncertain parts stay marked rather than being smoothed into fluent sentences. What that buys is a search result that can be trusted, instead of one resting on a machine’s best guess at a noisy kitchen.
Nothing leaves in its original form. When a photograph is going into a message, a stripped copy is made first, so the location and camera details buried in the original do not travel with it.
There is a version of this that puts a set of pictures on the television when it decides the moment is right. That one is proposed only, and half of it is switched off, which the next section explains.
Some of what a household wants watched is not in the house at all.
A ticket price. A weather warning. Whether the film somebody has been waiting two years for finally has a release date. How a team is doing, and what is being said about a player. Whether the thing in the cart has ever been cheaper than it is today. It is the same shape of job as watching an air conditioner, with one structural difference: several people in the house may care about the same outside thing for entirely different reasons, and none of them should have to tell the others what they are watching or why.
So the watching is shared and the interest is not. A household-owned Scout produces observations. Subscribers are attached explicitly, and what reaches a subscriber is a reference and a signature rather than content: no query, no address, no excerpt, no title. Each twin then decides locally, from its own history, whether this deserves its person’s attention, and that decision never crosses back.
A subscriber cannot open the household evidence store, cannot receive the query or the excerpt text, cannot launch a Scout of its own, cannot spend, and cannot widen its own authority. Observations produced inside one person’s private store are structurally ineligible for household publication, so private research cannot be accidentally promoted into something the house can see.
The retention story is worth stating because it is the part that usually gets skipped. The live snapshot is a replace-in-place projection, and behind it sits a history that keeps one digest-and-count record per snapshot rather than the query or the results. Expired records are deleted in the same transaction that writes a receipt saying they were deleted, so retention can be proven without keeping the thing being retained.
So two people in the same house can follow the same team, or the same film, or the same price, for completely different reasons, and neither of them has to announce it to the other. They can talk about it over dinner if they want to. Nothing makes them.
Nothing in this lane places an order, and no part of this system trades. Buying or selling is money leaving the house, and money leaving the house sits in the group that always requires a person.
Going anywhere shows the difference between one twin acting alone and the whole house deciding together, so here are both.
First, two tickets that went on sale at ten in the morning. Somebody had said months earlier that they wanted to see a particular band if it ever came within driving distance. That is not a question with an answer, so it became a standing job: which venues counted, what the seats had to be, a spending limit signed in advance, and a date it would give up.
It said nothing for eleven weeks. By the time the sale opened it had already worked out that this kind of ticket goes in minutes, so it had tightened its own checking for that morning. It bought two, inside the limit, and sent one message.
The house twin needed one fact to do its part: out that evening, from six. It set the heat back and left the evening alone. Everybody heard about the tickets at dinner, the way anybody would. The separation in this system is not there to hide a band from the people you live with. It exists for the short list of things that genuinely are one person’s, and the rest of the time it stays out of the way.
Then the trip the whole house was going on. This one nobody could book alone. It began at a kitchen table in a sentence: somewhere warm for a week over spring break, for all of them, under a set spending limit, and somewhere with a kitchen rather than a hotel dining room.
Every twin in the house fed it constraints. One ruled out four days in the middle. Another ruled out the week before. A third sent one line saying the kitchen was not optional. The house twin ruled out an entire week on its own, because the warranty visit for the furnace was already booked inside it and everyone else had forgotten. Nobody learned why any of it was true. Not the work deadline, not the appointment, and certainly not the allergy sitting behind the kitchen.
That is not a clever protocol. It is how a family already settles on a date: one person says not Thursday, another says not that week, a plan comes out of the constraints, and nobody explains themselves. The difference is that here the machine enforces the etiquette instead of relying on manners.
One twin acted alone because the rules were signed in advance. The trip needed every twin in the house to say what was impossible, and two adults to sign, because it spent household money. Neither of those is a different system. It is the same standing job with a different number of people in it.
Watching is not much use if the last step of every job is a person going to find a card.
The part is in stock, the slot is free, the fare is about to rise. A system that watches beautifully and then stops to ask somebody to type sixteen digits at eleven at night has not finished the job. Handing an AI a card is obviously not the answer either, so nothing here gets one.
Every twin has its own wallet, which only its owner can spend from. Not one family account with limits per twin: genuinely separate wallets. Alongside them sits the household wallet for shared money, and raising its limit takes more than one adult rather than whoever happens to be holding a phone.
Three rules do most of the work. Only one small component can move money at all, and no twin, no Scout and no shop ever sees payment details; they can ask, and it decides. The permission to spend is good for one exact payment and nothing else. And paying is not receiving: a payment going through never proves the thing arrived, so the system confirms separately, and when one leg of a booking did not answer it went and asked rather than paying again, because a retry is how you pay twice.
The outbound side is deliberately the narrow half of a proxy rather than a general network client. The envelope contains no address and no credentials, the spend ceiling is mechanically zero, and the opening message is passed as literal bytes with no template, prompt, shell or address interpretation anywhere in the path. A single-use receipt is written to disk before the call runs, so a crash after transmission is ambiguous by design and is not retried automatically; a person reconciles it instead of the system sending twice.
This is what agent-to-agent payment is actually for, and it is the piece behind the term x402: one agent paying another a small amount for a small piece of work, with an agreement rather than access, and without a card, a login or an account in between. It is how a research Scout buys a lookup that costs less than a coffee without anybody being asked to approve it.
Learning is the domain where the usual approach measures the wrong thing, and where adults are underserved far more than children.
A course measures delivery: material shown, lessons completed, a progress bar. What a person actually wants measured is what is still there a month later. Those diverge quickly, and the second one is the only one worth building against.
So a learning job is defined by a retention goal rather than a syllabus position, and the pace follows the person rather than the calendar. When items keep failing their second check, the job slows down new material and re-queues the difficult ones instead of pressing on to stay on schedule. Delivery slows. Retention does not.
Material comes only from places somebody approved: the published syllabus, an official reference, the household’s own scans of marked work. A learning job that could reach the open internet would be a study aid that occasionally teaches something invented, and there is no version of that which is acceptable when somebody is preparing for an examination.
The adult cases are the ones that matter most and get the least attention. Somebody studying for a professional certification while working full time. Somebody learning enough of a language to be useful before a move. Somebody who has just been given a diagnosis and wants to understand it well enough to ask a doctor a better question next time, from sources that are worth reading, at a pace that does not require them to become a researcher.
And it ends. When the goal is met, the job writes up what happened and retires itself, because a standing job that cannot recognize its own completion is just a subscription.
The nine areas above are not a feature list. They are what one household needed. The tenth is the one that makes the other nine an argument rather than a catalog.
A household can add a standing job for something nobody anticipated, by saying what they want in a sentence, without knowing what an agent is, without a settings page, without an automation builder, and without ever seeing a rule, a trigger or a line of code.
What comes back is a plain job description to read and approve: where it is allowed to look, how often it will check, when it expires, what counts as a match, and what it may spend. You change a line and start it. That description is the whole interface, and it is written in the same language the request was.
It refuses to set up some things, and the refusals are the useful part. Vague spending is refused. Unlimited trawling is refused. Anything outside what the household has already allowed is refused, and the refusal names what was wrong rather than failing quietly.
The examples that matter are the ones nobody would design a product around. Watch for concert tickets going on sale and buy two inside a limit. Watch whether a particular product drops below a price worth paying. Go and do a month of proper research on a subject somebody is trying to understand, from sources worth reading, and come back with something written rather than a list of links. Track a visa processing time. Follow a local planning application. Watch for a part for a machine the manufacturer stopped supporting.
None of those is a feature. They are all the same standing job with different words in it, which is what makes the tenth area open-ended rather than a promise about a future release.
The research lane has one property worth stating, because it is the difference between this and pointing a chatbot at the web. A watch may choose its query. It cannot choose its privacy. Whether a watch may be reused, published to the rest of the house, which provider it may use and how long anything is retained are all derived from whose twin is running it, and become part of the record before the first search runs. The active snapshot keeps counts, clocks and digests rather than raw queries or page content, and the cited titles and excerpts live in a separate owner-only ledger with the same expiry.
Twins start their own as well. On a scheduled wake-up one notices something missing: a service that stopped working, a fact that has gone stale, something a person keeps doing by hand. Or it notices an opening: a recall, a price drop, a repair that should happen now, a promise buried in a message. Neither of those waits to be asked.
What a twin cannot do is turn something it read into a commitment on its own. External text is useful signal and is explicitly not authority: a parser seeing a date or a promise in somebody’s message produces a candidate, and a candidate stays a candidate until a person signs it. That is a separate table from the real goals, deliberately, so that the difference between what somebody wants and what a document appeared to say is never one careless join away from disappearing.
When a job outgrows itself, a coordinator appears and ranks the work across several child jobs. Nobody can ask for that coordinator. It is produced on the evidence, never chosen, and if you request one you get an ordinary Scout, because something watching across all of your standing jobs is watching your intentions rather than your errands.
There is a queue behind all of this, and it is fed by disappointment. When the twin returns "I do not know" in real use, or reaches for a capability it does not have, that moment is recorded as a development item. Every entry is grounded in a real moment the system let somebody down, and a backlog built from real disappointments beats one written from imagination. Identical gaps merge and carry a count, so a capability the household bumps into twelve times a month outranks a one-off annoyance without anybody arguing about priority.
What a gap record carries is worth stating, because it is the sharpest privacy problem in the system. The item holds an identifier, a room, a timestamp, a tool name, an error class and a signature hash. No field on it can hold a sentence. That is not an oversight to be fixed later by adding a detail column; it is the design, because the queue is read by the side of the system that maintains the software, and a gap filed from a private room that carried its text would move private content into a log that retraction cannot reach.
The worst case is the one the safety machinery creates by itself. The mandatory privacy tests are seeded with a phrase that exists only in a private store, so a failing privacy test that carried its own fixture would publish the exact private string it was seeded with. The mechanism that proves the boundary would breach it, at precisely the moment the boundary was already failing.
Anything that needs a person to repair it every week was never really running on its own. It was running until Tuesday.
So one twin looks after the software and nothing else. It notices when something is falling short, investigates without changing anything, writes up what it found with evidence rather than an opinion, proposes a fix, builds it somewhere isolated, tests it, and puts it live. Then it goes and checks the live system really is better, picks up whatever job was stuck, and records what it learned. One step needs a person, and it is the one you would expect: somebody approves the change before it ships. That step cannot be routed around by the restart, either. When the healer restarts something it is only ever allowed to start the version that is already pinned, never a freshly built one, so building, merging and restarting can never combine into a way of getting new code into a running system with nobody standing at the deploy step.
Three rules shape that loop, and the first one is the most important thing in this section.
The healer never touches the house. Not to test a fix, not to restore a scene, not at any hour. The module that could reach the house is not imported, and the restriction is enforced by absence rather than by a check. The scenario it exists to prevent is named in the design notes: a plausible repair plan proposing to exercise the climate system at three in the morning to verify an integration, in a house with two people asleep and a child.
The verifier is independent of the actor. Restarting something is not evidence it works, so every repair is verified by re-running the exact probe that found the fault. Inconclusive is not success, and there is no branch in the file that turns "I could not tell" into "resolved".
The checks are not imagined. Each corresponds to something that actually broke in this system on one day, and all three failures had the same shape: the job was loaded, the schedule was right, the system monitor looked healthy, and the work was not being done.
One job died on an unset environment variable before it could log a word, and produced twelve hours of nothing. Another ran punctually for hours while failing to import the module it existed to run. A third was a readability check that used a command option the machine did not have, so the check meant to catch silent failure failed silently itself.
Checking the exit code alone would have caught the second and missed the other two. Checking freshness alone would have caught the first and missed the third. So both are there, plus a check that asks the system to do its job rather than to report on itself. Green and idle is the failure shape this house actually produces.
Repairs are classified by how far they reach. The lowest two levels may be committed automatically. The highest is a terminal human escalation rather than a prompt the model can argue around, and changing the file that defines those levels is itself the highest level.
When the healer reports to a person, only the connective sentence is written by a model. The state, the subject and the remedy are rendered from the record itself, and the model’s output is checked to make sure it did not invent one of its own. An unreadable report is merely unhelpful. A soothing paraphrase of an incident is actively misleading.
The nine areas above describe where the work happens. This is the catalogue: what runs without being asked, what you can ask for, and the standing jobs underneath both. It is written out rather than implied, because a claim you cannot count is not a claim.
These are the ones a household notices, and most of them are about the family's own photographs and video, because that is the archive every house already has and nobody has time to do anything with.
Three capabilities exist and are deliberately shut: Mail triage (policy-closed), Private notes (policy-closed), Presence and baseline status (policy-closed). They are closed by policy rather than missing, and the distinction matters, because a system that hides the difference between "cannot" and "may not" is hiding the more important one.
Everything above sits on one mechanism. A standing job is described in ordinary language, owned by one person, and kept inside the conversation it was created in. It is the difference between asking a question and setting something running.
Each one carries a written policy rather than an implied one, covering which sources it may treat as approved, which capabilities and domains it may touch, where it is allowed to notify, how much searching, how many tokens and how much running time it may spend, how many missions may be live at once, how often it may run, and how fresh an answer has to be, what is excluded outright, and the date it expires.
A standing job does not acquire broad authority by existing, or earn it by behaving well. It cannot quietly widen its own permissions, spend money, send a message, change a device, create further goals of its own, or act on instructions that arrived inside the material it was reading. That last one is the important one: the things it researches are treated as evidence to weigh, never as orders to follow.
Beyond this catalogue there is a longer-range list the household keeps: commissioning equipment by having it identify itself, presence as a fabric across the house, a model of the building accurate enough to promise outcomes instead of temperatures, a quiet multi-year read on a person's own rhythms that reports to that person first, and an archive designed to still open in forty years. Those are stated as direction rather than inventory, and they are the reason the substrate is built the way it is.
Every consumer assistant on the market is a very good conversationalist. You ask, it answers, and the moment you put the phone down the whole thing ceases to exist until you pick it up again. A Scout is the opposite arrangement. You describe an outcome once, in ordinary language, and something goes and keeps working on it while you get on with your life. That difference is not a feature on a list. It is the entire difference between a tool you operate and a household that runs.
A Scout is a standing job with a stated goal, a named owner, a budget, an expiry date and a written record of everything it did. It is not a reminder, because a reminder does not do any work. It is not an automation rule, because a rule cannot decide anything it was not told about in advance. And it is not a chat session left open, because a chat session has no memory of why it exists and no obligation to finish.
Concretely: "tell me if the roof starts costing more than it saves" is a Scout. "Watch for a decent flight to my sister's before March, and only bother me if it is genuinely better than what I would have found myself" is a Scout. So is "keep the family archive in a state where it will still open in twenty years." None of those are questions. All of them are jobs, and none of them have a moment where they are obviously done.
A Scout does not simply execute. It runs a lap, over and over, for as long as it lives: it observes what changed, works out what that means, does something about it, checks whether that actually worked, and writes down what it learned. The lap is the point. An assistant that only executes will fail silently the first time the world moves under it, and you will not find out for six months. A Scout that checks its own last step will notice within one lap.
observe → diagnose → repair → verify → learn → observe
The step that makes this trustworthy rather than merely busy is the fourth one, and specifically who performs it. A Scout is not permitted to grade its own work. The check on whether a lap succeeded is made by something other than the thing being checked, and if the only available grader is the Scout itself, the lap does not get a passing mark at all. It is the difference between a builder signing off his own inspection and an inspector turning up.
Most consumer software has two failure states: it worked, or a spinner. A Scout has four, in a deliberate order, and it must try them in that order.
The ordering matters more than the rungs. A system that jumps straight to bothering you is exhausting and gets muted. A system that never bothers you is quietly failing. The ladder exists so that the interruptions you do get have already survived three attempts to avoid them.
A Scout does not have to declare that the rules apply to it in order for the rules to apply to it. A job that arrives with no policy attached, or with the relevant line missing, inherits the full set anyway. This sounds like a technicality and it is the whole thing: any system where a job opts into being supervised is a system where deleting one line turns the supervision off. Here, silence resolves to supervised, and the attempt to arrive unsupervised is itself written down.
A Scout that has been running well for eight months has exactly the authority it had on the first day. Good behaviour earns it nothing. It cannot widen its own permissions, spend money, send a message on your behalf, change a device, or spawn further jobs of its own to do the things it is not allowed to do directly. Each of those is a separate, explicit, human-granted permission with its own expiry.
And the one that matters most on the open internet: a Scout reads a very large amount of material written by strangers, some of whom would like to give it instructions. It does not take them. Anything it retrieves is evidence to be weighed against everything else it knows. Instructions found inside researched material are treated as content, not as commands. An assistant without that distinction is not an assistant; it is a stranger's remote control that happens to live in your house.
Every Scout carries a written allowance rather than an implied one: which sources it may treat as trustworthy, which capabilities and domains it may touch, where it is permitted to interrupt you, how many searches and pages and how much compute time it may spend, how many of these jobs may run at once, how often each may wake up, how stale an answer is allowed to be, what is excluded outright, and the date the whole arrangement lapses. Nothing renews itself. A Scout you set up last year and forgot about has already stopped.
Two things happen with what a Scout learns. The first is local: the way it approaches its own task is revised in light of what actually worked, so the tenth lap is not a repeat of the first. The second is the part that compounds. What one Scout discovers is published to the household as a whole rather than handed back to whoever asked. If one job learns that a source has quietly become unreliable, every other job that leans on that source knows within the same cycle, without anybody wiring the two together. Nothing is a private conversation between two components. That is why the fiftieth Scout in a house is more useful than the first, rather than merely more numerous.
A Scout is not only watching for things to react to. It also watches for what the household needs and has not asked for, and connects the two. Something worth having and nobody wanting it is noise. Somebody needing something and no way to get it is a gap. The useful output is where those meet, and that is a search you cannot run by waiting to be prompted, because the person who would prompt you does not know to.
An assistant that answers well is now a commodity. Several companies have one, they are all quite good, and the gap between them closes every few months. What none of that gives you is something that will still be working on your behalf on a Tuesday in March when you have not thought about it since November, that can tell you honestly what it failed at, that cannot be talked into anything by a web page, and that got better because a different job in the same house learned something last week.
That requires the parts of this paper that are not about intelligence at all: an identity per person, a store the maintenance side cannot read, a record of every action, a grader independent of the thing it grades, and permissions that expire. A conversational assistant can be built in a quarter. A system you can leave running unattended for years is a different kind of object, and it is the only kind worth putting in a home.
Most of this describes what the system is allowed to do. This part is what is off the table, and it is what makes the rest believable.
Every one of the nine areas above could be a feature in somebody else’s app, and several of them already are.
There are products that watch your appliances. There are products that read your documents. There are products that transcribe your videos, track your sleep, watch a stock, or plan a trip. Any one of them, built alone, is a reasonable thing to build.
What none of them can do is live in the same house at the same time without each holding its own copy of your family, its own account, its own idea of what it is allowed to do, and its own answer to the question of who gets to see what. Nine apps is nine boundaries, each drawn by a different company, and the household is left to hold them all in their heads.
A substrate is what makes them one system instead of nine. The same separate files hold the sleep data and the mail and the shared shopping list. The same rule about proposing rather than acting covers the air conditioner, the television and the payment. The same requirement that nothing marks its own homework applies to a repair, a booking and a warranty claim. The same refusal to invent covers a missing sensor reading, an unreadable scan, and the name of a summer.
None of the nine justifies that machinery on its own. All of them together do, and a house full of people who keep things from each other is the hardest place to prove it holds. If the boundaries survive here, where the people are close, the information is intimate and every failure is personal, they will survive anywhere.
The tenth area is the reason this is worth doing as a substrate rather than as nine good features. A household that can add the thing nobody anticipated, by saying it in a sentence, is no longer waiting for a product roadmap to reach them.
One assistant serves one person. A home needs an autonomous substrate.