hindsight, engineered

AI that answers the question you asked.
Not the one it thought sounded better.

Hindsense is a set of tested guardrails for the professionals who use AI to research, draft and deliver real work (consultants, engineers, scientists, analysts, planners). It turns a plausible-sounding assistant (the one you already use) into a partner that returns results you can act on, sign, and defend.

  • Interrogates every prompt so the AI targets your actual objective, not a nearby one
  • Holds the draft to a standard you would apply to your own work, not the AI's default
  • Refuses fabricated citations, invented clauses and guessed names before they land in a deliverable

You keep your assistant. Hindsense sits behind it.

Keep using Claude, Copilot, ChatGPT, Perplexity, Gemini, whatever you already trust. Hindsense is not another chat window. It is a guardrail suite that plugs into the assistant you already use, runs in the background, and only speaks up when the work is drifting off-objective, thinly supported, or about to land a mistake in your deliverable. Deploy it once and work the way you already do.

  • Assistant Bring your own. Hindsense sits alongside it, not in place of it.
  • Trigger Runs when you ask, or when the assistant tries to hand you a finished draft.
  • Output A verdict and a reason, back into the same conversation. No new window, no re-paste.

The AI gave you an answer. It just is not an answer you can use.

The failure that costs you time is not the answer that reads wrong. It is the one that reads right, sounds authoritative, arrives fast, and is quietly off-target (the wrong question answered well, the wrong assumption baked in, a fabricated reference at the end). Discovered by your reviewer, your auditor, your regulator, or the person on the other side of the table.

off-target

The nearby question

You asked about the risk to a specific receptor under a specific set of conditions. The AI answered a more general version of the question that sounds thorough and misses the point. You only notice on the third read.

unchecked

The confident guess

The AI states a limit, a threshold, a clause number, a person's role or a company's address with no hesitation. None of it is verified. Some of it is invented. Everything about the draft says trust me.

shallow

The default-safe conclusion

The AI hedges to the most conservative reading, because that is what its training rewards. On your work, a needlessly conservative conclusion is a defect, not a safety margin (you have to justify it either way).

brittle

The draft you cannot defend

The prose is competent. Ask what a specific paragraph rests on and it dissolves. No source, no reasoning trail, nothing to hand to a reviewer or an auditor. You end up rewriting from scratch to make it defendable.

A layer of professional discipline the AI does not have on its own.

Every guardrail below started as a mistake that survived a review in real work, was diagnosed, tested against dozens of jobs, and codified so the next assistant does not repeat it. Two of them do the heaviest lifting (one at the front of the work, one at the end), and everything else clips into that spine.

the front of the work

Prompt refinement

Before any research or drafting starts, the guardrail forces a real objective onto the request. What are you actually trying to decide? What would a right answer look like? What is in scope and what is not? What are the failure modes for this kind of question? A vague prompt returns vague work, so the vague prompt gets sharpened before it costs you a wrong deliverable.

  • Names the objective, the success test and the scope up front
  • Surfaces the assumptions the AI would otherwise smuggle in
  • Rejects the prompt if it cannot answer what a right result would look like
  • Stops research fanning out into ground you never asked about

the end of the work

High-quality outputs

Once the draft exists, the guardrail runs it against the standards you would apply to your own work. Not a grammar pass. A first-principles verification: are the numbers internally consistent, do the assumptions actually hold, is this the best supported conclusion or the safest-sounding one, and are there specific reasoning trip-hazards this kind of question is known to fall into? Findings come back for your decision, not silent rewrites.

  • Verifies factual claims from first principles rather than pattern-matching
  • Treats an unnecessarily conservative conclusion as a defect, not a safety margin
  • Names the assumptions a claim rests on and whether they hold
  • Flags the specific reasoning traps that catch this kind of work

Between and around those two, a full suite of guardrails handles the mechanics (the references, the tables, the files, the memory of what was decided last session). The citation-discipline engine is live today; the rest of the suite is what the waitlist is for.

Every guardrail is a mistake we saw once and never wanted to see again.

One line each. Written for somebody who has never seen the product. The citation-discipline engine is live today, its verdicts run against real drafts. The rest are specified, tested against real work, and being wrapped into the product (which is why this is a waitlist and not a shop).

Every deliverable needs these

Objective and evidence

  • Sharpens a vague prompt into a real objective before the AI runs on it
  • Refuses a citation whose source could not hold that kind of claim live
  • Blocks a standard, statute or guideline reference with nothing behind it live
  • Refuses to fill in a name, role or contact detail that was never confirmed live
  • Pushes you to the original source rather than the article that summarises it
  • Checks units, datums and reference conditions match before a number is compared to a limit
  • Re-derives a pass or fail rather than trusting the one already typed into the table

Drafting that reads like your work

  • Answers the question first, then supports it, instead of building slowly to a conclusion
  • Verifies factual claims from first principles rather than pattern-matching a familiar answer
  • Treats an unnecessarily cautious conclusion as a defect, not a safety margin
  • Raises a substantive problem for your decision rather than quietly rewriting what you meant
  • Catches the same passage repeated in two sections, and content sitting in the wrong section
  • Tests that every commitment in a plan names a trigger, an action, an owner, a timeframe and a record
  • Writes to somebody who already knows their field, rather than explaining it back to them

Files that survive the round-trip

  • Checks that a Word file actually opens in Word, not only in the tool that produced it
  • Repairs cross-references so section, table and figure numbers update instead of breaking
  • Keeps a spreadsheet's inputs, workings and results separate, and its units consistent
  • Assembles a photo record with captions that describe what is there, and nothing more

A project record that holds up

  • Holds every document you supplied, so you are never asked to send the same one twice
  • Treats the version you edited and sent back as the one that governs from then on
  • Records what was decided so a later session cannot quietly reverse it
  • Keeps ideas still being explored separate from conclusions you can act on
  • Distinguishes what has been built from what is actually live
  • Produces the AI-declaration record your quality system asks for at sign-off

Discipline-specific, in development

The first specialist pack, drawn from environmental consulting. Later packs will follow for adjacent disciplines. If your work is not represented, the waitlist question is the place to say so.

  • Reviews underground service and utility plan responses before intrusive works, so a live asset does not get hit
  • Drafts the regional geology and hydrogeology sections a site owner or contractor can actually visualise
  • Interprets soil and groundwater results against the assessment criteria the report is being written to
  • Assembles a site inspection photo record with what-is-there captions, not sampling-equipment commentary

You get a verdict and a reason, not a confidence score.

Every finding names the claim it belongs to, its status, its severity, and the evidence that was actually supplied. Below is real output from the live citation-discipline engine, not an illustration. The other guardrails return findings in the same shape.

POST /v1/skills/citation-discipline/invoke
"verdict": "refuse",
"findings": [
  {
    "target_id": "c1",
    "status":    "verified",
    "severity":  "info",
    "message":   "Verified against provided source structure.
                 Reconfirm the exact clause number in the
                 in-force version before signing."
  },
  {
    "target_id": "c2",
    "status":    "unsupported",
    "severity":  "blocker",
    "message":   "No source supplied for a clause-level claim.
                 Preferred publishers: state or commonwealth
                 legislation registers, the standards body
                 itself, or the regulator's own guidance page."
  },
  {
    "target_id": "c3",
    "status":    "named-but-flagged",
    "severity":  "warning",
    "message":   "Source is present but not from the preferred
                 publisher for this claim kind."
  },
  {
    "target_id": "s1",
    "status":    "named-but-flagged",
    "severity":  "warning",
    "message":   "Internal specific supplied without a verification
                 channel. Do not expand initials, usernames, or
                 email local-parts to a full name unless verified."
  }
],
"entitlement": { "ruleset_version": "2026.08", "update_available": false }

3

verdicts (pass, pass-with-flags, refuse). No probability to interpret.

5

finding statuses, so a flag can be triaged rather than dismissed.

1

guardrail live today. The rest are the suite above, wrapping into the product now.

What people ask before signing up.

Does AI actually make things up in professional reports?

Yes. Language models generate plausible-looking citations, standards references, clause numbers, dates, addresses and names that do not exist, and they do so with confident wording. The failure is systematic, not a bug awaiting a fix. Which is why sitting a guardrail between the assistant and your deliverable is the fix, not switching models.

Is this different from just asking the AI to be careful?

Yes. An assistant told to "always cite your sources" will happily cite a fabricated one. Told to "think carefully" it will produce longer prose in the same shape. The Hindsense guardrails are external checks with their own logic. They interrogate the prompt before the AI runs, and evaluate the draft against tested standards after, rather than trusting the AI to police itself.

Which AI does it work with?

Any assistant that can call an external skill file. The engine is a JSON API (the assistant sends the material to check, the engine returns findings). The same guardrails have been used against several frontier models in real work.

Is this shipping today?

The guardrails exist, they have been tested against dozens of real jobs, and they work. The citation-discipline engine is live behind a stable JSON contract. What is being built is the product wrapper (accounts, keys, the way the rest of the suite is packaged and distributed). That is what the waitlist is for.

Who is it for?

Professionals whose work is judged on whether the outcome is right and defendable (consultants, engineers, scientists, analysts, planners, technical specialists). If you use AI to research, draft or review, and a wrong answer costs you real time or real credibility, the waitlist is aimed at you.

What happens to my draft when the engine checks it?

The engine evaluates the material in the request and does not retain the draft after the response. Waitlist signups store your email address for one release notification, and nothing else. No marketing platform, no third party, no sale.

A product about substantiation should be honest about its own limits.

  • It does not replace your judgement. The guardrails surface findings, trade-offs and specific defects for your decision. They do not sign the deliverable, and no verdict is a substitute for your review.
  • It does not fetch a source and read the clause end to end. The live engine checks that the source you supplied is structurally plausible and is the right kind of publisher for the claim. Deeper live retrieval is a later capability; the wording of every finding is careful not to imply otherwise.
  • It does not guess at people. Where a name, role or contact detail cannot be tied to a verification channel, it stays flagged rather than being quietly resolved into somebody's full name.
  • The product wrapper is not shipping yet. The guardrails are built and tested. Accounts, keys and the distribution of the wider suite are in development. There is no release date to promise you, so there isn't one on this page.

Hindsight, before you need it.

One note when the first release is ready. Nothing before then.

Join the waitlist