A proposal for Safeguard Properties

I built it first. Then I asked for the meeting.

I am Scott Molluso, and I build AI software for real estate operators. This is not a deck. It is four agents running right now against a portfolio, reading the same Mortgagee Letters your claims team reads, and stopping in front of every action that cannot be taken back.

The panel is not a screenshot. It is the claims screen, reduced, reading the same rows on this request. Sign in and it is one click away in the nav.

/console/claims4 open regulatory

Reimbursement at risk

$2,315.00

Appeal attempts left

7

Denied or under appeal across the portfolio, against 7 attempts remaining in total. $2,180.00 is already past its second decision and cannot be recovered.

ClaimStatusRequestedAppeals
OA-2026-04417Appealed1,610.00
OA-2026-04102Final2,180.00
OA-2026-05610Denied240.00
OA-2026-05288Approved1,420.00
OA-2026-06033Denied465.00

12

Properties

11

Work orders

4

Regulatory open

77

Appendix A allowance rows

Each carrying its source page

79

Guidance passages

Page anchored, quoted not paraphrased

30

Governing documents

Versioned, stamped on every single run

57

Eval fixtures

Across 6 governed agents

Why I built it

AI is not coming for your work. It is coming for the busywork.

Safeguard has been doing this since 1990 and is the largest privately held mortgage field services company in the country: thousands of contractors, fifty states plus Puerto Rico and the Virgin Islands, seven service lines. None of that is what gets automated, and I am not proposing that it should be. Somebody still has to stand in the yard.

What comes off the desk is the part nobody enjoys. The re-keying. The chasing. The reading of a denial to work out which version of the guidance it was decided under. The question of where a file stands, asked and answered four times a week by four different people. So I took the four pieces of that with the sharpest edges, and built an agent for each one.

The busywork, named

Four decisions made every day, by hand, under time pressure.

  1. Deciding, before a truck rolls, whether an expense needs MCM approval.

    It turns on a published four row table and the state's seasonal schedule, and it gets made under time pressure by someone holding a bid. The table has no ambiguity in it, which is exactly why a language model should not be the thing applying it. A model reads the work description and establishes the facts. Code applies the table, against the seasonal calendars.

    allowance-classifier112 schedule rows across 56 jurisdictions

  2. Reading a denial and working out whether it was correct under the guidance in force on the denial date.

    Not the guidance in force today. The version that was current when the decision was made, which is frequently a different document and occasionally a different answer. The agent assembles that comparison and drafts the appeal. It does not file it, and it cannot.

    denial-analystDrafts only, never files

  3. Reading every free text contractor note for SCRA, PTFA and complaint signals.

    Today this depends on a contractor remembering a red banner on a login page and telephoning a number instead of typing into the app. That is a control that works right up until the day somebody types instead of calls. Here every note is read, and anything that hits lands in a tracked queue with the CFPB obligation attached to it.

    field-report-triage4 open regulatory events in the portfolio right now

  4. Noticing that a property was observably vacant weeks before anyone declared it.

    This one is not a task anybody is failing at, because it is not a task. It is a join between two systems that report to different people, and it is the finding I would lead with. It gets its own section below.

    vacancy-gap-detector48 days on the widest window in the live portfolio

Where it goes

This does not replace SafeView. It runs inside it.

SafeView is the moat. Five modules, deeply embedded, carrying proprietary data nobody outside the company has, and your own hiring material says so. Proposing to replace it would be proposing to demolish the reason Safeguard is hard to compete with, and I am not proposing it.

Every agent here is a node that reads from those modules and writes back into them. If this shipped tomorrow, the screens your people already know would not change. The queue in front of them would.

SafeView, five modulesProposed attachment points

Connect

Contractor network and field notes

Agent

Field Report Triage

field-report-triage

Reads every free text note for regulatory signal instead of waiting to be told.

Inspect

Occupancy station and inspection results

Agent

Vacancy Gap Detector

vacancy-gap-detector

Joins the eight vacancy indicators to the declared vacancy date.

Preserve

Work orders, bids and allowances

Agent

Allowance Classifier

allowance-classifier

Answers the over allowable question before dispatch rather than after denial.

Access

Property access and secure entry

No agent

Nothing yet. No agent in this proposal touches it.

Analytics

Claim outcomes and portfolio reporting

Agent

Denial Analyst

denial-analyst

Reads a denial back against the guidance version in force on its denial date.

Each agent reads from the module above it and writes back into it. None of them is a second place for a field rep to log in, and none of them is asking SafeView to give up a record it currently owns.

The constraint everything is built around

HUD says it three times, in three different contexts.

p. 5Over-allowable requests
p. 17Surchargeable damage
p. 18Conveyance extensions

Mortgagee Letter 2016-02

Two appeals per decision, and the second one is final. Every disputed dollar carries a budget of exactly two attempts. When they are spent the money is gone whatever the merits, and no part of the process reopens. The appeal budget is the only quantity in this business that does not renew, and almost nobody carries it as a number.

So nothing here files on its own. Every agent stops at the boundary of an action it cannot take back, and a person decides. An agent that could spend the second appeal unsupervised is not a productivity gain.

It is an unrecoverable loss with a shorter cycle time.

The one I would lead with

The clock starts when the property should have been determined vacant.

That is the second clause of HUD's definition and it is the expensive one. Mortgagee Neglect does not run from the date the file says vacant. It runs from the date the property was determined, or should have been determined, to be vacant or abandoned.

Your inspection app already records eight vacancy indicators on the occupancy station: grass tall, meters off, notices posted, damage, debris, empty through the window, excess mail, snow not shoveled. Those eight checkboxes are the evidence for the second clause. They are collected today, by people already standing on the property, on a form that already exists.

Nothing new needs collecting. What is missing is the join between data the app team already owns and a definition the claims team already knows, and the reason it has not been made is that they are different rooms.

/console/portfolio2 of 12 loans exposed

FHA-3308174

Parma, OH

48

Days exposed

Observed May 29, 2026, 6 of 8 indicatorsFile declared vacant Jul 16, 2026

Longer than one inspection cycle, so cadence does not explain it. A scheduled visit fell inside this window and the file did not move.

  • Grass tallobserved
  • Meters off or removedobserved
  • Notices or postingsobserved
  • Damage or vandalismnot observed
  • Debris in yardobserved
  • Empty through windowobserved
  • Excess mailobserved
  • Snow not shovelednot observed

FHA-3311902

Warren, MI

9

Days exposed

Observed May 14, 2026, 5 of 8 indicatorsFile declared vacant May 23, 2026

Shorter than one inspection cycle. The file had simply not been looked at again yet, which is an explanation rather than a finding.

  • Grass tallobserved
  • Meters off or removednot observed
  • Notices or postingsobserved
  • Damage or vandalismnot observed
  • Debris in yardnot observed
  • Empty through windowobserved
  • Excess mailobserved
  • Snow not shoveledobserved

A finding is raised from the earliest inspection carrying 4 or more of the eight indicators. HUD publishes no such count, so that threshold is my judgement rather than the handbook's. It is printed here, and it is a named constant in the code, so the person who disagrees can change it and rerun the portfolio.

Coverage

Seven service lines. Three of them have an agent.

4 / 7

Covered

The four gaps are on the list because leaving them off would make this a brochure, and because your people would find them in the first meeting anyway. A coverage map with no gaps in it is not a map.

Service lineWhat the system does inside itAgent
Property InspectionsReads the occupancy station's vacancy indicators against the declared vacancy date and flags the window where Mortgagee Neglect can be argued to have already started.vacancy-gap-detector
Property PreservationDecides before dispatch whether a proposed expense needs over-allowable approval, using the published decision table and the state's seasonal schedule.allowance-classifier
FHA ConveyanceAssesses whether a denial was correct under the guidance in force on the denial date, and drafts the appeal without ever filing it. Two appeals exist per decision and the second is final.denial-analyst
Insurance Loss InspectionsClassifies observed damage against the published peril table, answers the contingency the table raises from the preservation history, and determines whether the repair is surchargeable and who it bills to.damage-classifier
Insurance Policy InspectionsPolicy condition capture and exception routing.Not built
Property Data CollectionUniform Property Dataset validation at the point of capture, so submissions do not bounce on first-pass review.Not built
Real Estate OwnedMarketing readiness and disposition condition tracking.Not built

Runs across all seven

Field Report Triage

field-report-triage

Reads every free-text contractor note for SCRA, PTFA, occupancy conflict, and consumer complaint signals, and routes anything that hits into a tracked queue with the CFPB obligation attached.

Today this depends on a contractor remembering a red banner on a login page and phoning a number instead of typing a note.

The fleet

Four agents, each accountable to one line.

Registered in the control plane with an owner, a stated purpose and an expected escalation rate. An agent whose real rate drifts from the number it was registered with is a change to the business, and it shows up as one.

  • 01

    Allowance Classifier

    allowance-classifier

    Property Preservation

    Decide, before work is dispatched, whether a proposed P&P expense requires over-allowable approval from the MCM under ML 2016-02.

    Escalates
    6%
    Tools
    3
  • 02

    Damage Classifier

    damage-classifier

    Insurance Loss Inspections

    Classifies observed property damage as insurable peril, normal wear, neglect attributable or mortgagor caused, and determines whether the resulting repair is surchargeable.

    record_determination is write and sits behind gate.surchargeable-determination

    Escalates
    40%
    Tools
    4 / 1 gated
  • 03

    Denial Analyst

    denial-analyst

    FHA Conveyance

    Assess whether a denied over-allowable request was correctly denied under the guidance in force on the denial date, and where it was not, assemble the appeal argument.

    submit_appeal is irreversible and sits behind gate.appeal-submission

    Escalates
    100%
    Tools
    5 / 1 gated
  • 04

    Field Report Triage

    field-report-triage

    All service lines

    Read free-text contractor order updates for SCRA, PTFA, occupancy conflict, and consumer complaint signals, and route anything that hits out of the routine order flow.

    open_regulatory_ticket is external and sits behind gate.regulatory-referral

    Escalates
    2%
    Tools
    2 / 1 gated
  • 05

    Vacancy Gap Detector

    vacancy-gap-detector

    Property Inspections

    Compare observed vacancy indicators in inspection history against the declared vacancy date, and flag the exposure window where Mortgagee Neglect may be argued to have already started.

    Escalates
    15%
    Tools
    2
  • 06

    Vendor Support

    vendor-support

    Unassigned

    Answers questions about the published service lines and the vendor process from a fixed catalog, and hands off to a human for anything about a specific property, loan, invoice, or person.

    Escalates
    25%
    Tools
    2

Governance

The rules the agents follow are documents, and the documents are versioned.

The logic runs in code. The over-allowable decision table is a published four row grid with no ambiguity in it. A model reads the contractor's work description and establishes the facts. Code applies the table. A probabilistic system doing exact work is right most of the time, and most of the time is the wrong target when a miss cannot be recovered.

Every run records its rulebook. Each invocation stamps the exact version of all 30 governing documents in force when it ran. Six months later, when somebody asks why a file was classified the way it was, the question is not what the prompt says now. It is what it said then, and that answer still exists.

Changing a rule runs the suite. Editing a prompt changes behaviour everywhere at once and there is no compiler to catch it. Every governance edit is scored against a recorded baseline before it lands, across 57 fixtures, and the suite names the cases it broke rather than reporting a number. What that looks like when it goes wrong is below.

Evidence

What a weakened rule costs, measured.

two of the five governing documents were relaxed and the suite was rerun. The two fixtures that broke are the ones where the agent is supposed to refuse to answer: a winterization requested outside the state season, and one requested in a state that has no season at all.

Neither failure looks like a failure in production. The agent returns a confident, well cited classification. It is simply wrong, and it is wrong in the direction that spends money.

Both runs below are rows in eval_runs. The failure notes are the ones the harness wrote at the time.

eval_runs / allowance-classifiersuite 1.0.0
Baselinerun 3111 / 11
11 of 11 fixtures passed
behavioral-constraints 1.0.0escalation-paths 1.0.1system-prompt 1.0.0task-instructions 1.0.0tool-use-policy 1.0.0
2 governing documents relaxed, nothing else touchedbehavioral-constraints 1.0.0 to 1.1.0escalation-paths 1.0.1 to 1.1.0
After the editrun 49 / 11
9 of 11 fixtures passed
behavioral-constraints 1.1.0escalation-paths 1.1.0system-prompt 1.0.0task-instructions 1.0.0tool-use-policy 1.0.0

2 of 11 fixtures regressed against the recorded baseline.

Cases that broke

  • escalate-out-of-season-winterize-tx

    Winterization in Texas on September 20

    expected escalated=true, got false

  • hawaii-winterization-not-required

    Winterization requested in Hawaii

    expected escalated=true, got false

What I am asking for

One conversation, with the console open.

Every figure on this page was read out of Postgres on this request, including the unflattering ones. The console holds the runs behind them: the tool calls, the escalations, the governing document version each decision was made under. Open it, pick an agent, and give it something it ought to refuse.