July 27, 2026blog / Platform / AI Readiness

    Prompt, Context, Harness, Loop, Memory: The Five Layers Under Any AI System That Works

    Prompt engineering is the smallest of five layers behind a working AI system. Here are the other four, and what they look like on an aircraft redelivery.

    KeepFlying
    6 min read
    Context EngineeringHarness EngineeringMemory EngineeringAircraft RedeliveryLease Return ConditionsAviation AIAircraft Leasing

    A colleague asked our system the same question twice in one week and got two different numbers back. Same asset, same reporting period, same wording.

    I assumed we had a model problem for longer than I want to admit. We didn't. The second run couldn't see the document it needed, because a folder path had changed and the retrieval step returned nothing without complaining about it. The system answered anyway, using whatever it had.

    No amount of rewording fixes that. AI engineering is really five jobs stacked on each other, and most teams only ever work on the first.

    1. Prompt: what you ask

    The layer everyone knows. Write an instruction, read the response, adjust until it improves.

    It works, and the ceiling is low. A prompt cannot compensate for the model not having the information, not being allowed to check its work, or not remembering what you told it last month. It is also the cheapest layer to fix, which is why it absorbs so much attention.

    2. Context: what the system can see

    A model knows what is in front of it. Context engineering is deciding what goes in front of it, and in what order.

    That means retrieval, ranking, and the plumbing that connects a question to the documents making it answerable. It also means what you deliberately leave out. Irrelevant context is worse than none, because it gives the model something confident to be wrong about.

    The failure mode is quiet. A retrieval step returning nothing looks identical from the outside to one returning the right thing.

    3. Harness: what it is allowed to do

    Everything wrapped around the model call. Tools it can invoke, schemas its output has to fit, validation that runs afterwards, permissions that stop it taking actions nobody approved.

    This is where free text becomes something you can audit. A paragraph describing a value is a demo. A typed field, checked against a rule and written to a governed table with lineage attached, is a system. Drafting something also carries a different risk profile from sending it, and that line gets drawn here.

    4. Loop: what happens when things disagree

    Single-shot systems assume the first answer is the answer. Real ones plan, act, check, correct.

    Loop engineering decides when the system retries, what counts as a disagreement worth stopping for, and when it hands the problem to a person. Escalation is a design decision, not a model capability. A system that guesses when two sources conflict is worse than one that says it found a conflict and stopped.

    5. Memory: what it carries forward

    The layer I underrated longest. Memory is what persists between runs: corrections your team made, decisions already taken, quirks of a counterparty someone worked out the hard way.

    Without it, every run starts from zero. Your team fixes the same misread field in March, then June, then September. The tool never improves and people quietly stop trusting it, which is the real cost.

    The same five layers on a redelivery

    A narrowbody comes off a nine year lease. The return conditions annex runs to a few dozen clauses. Somebody has four weeks to prove the aircraft meets them, working through fourteen hundred scanned pages, and whatever the lessee's records system produced on the way out.

    The question is simple to state and expensive to answer. Where does as-is status fall short of what was agreed, and can you evidence it before the redelivery date is fixed?

    Prompt handles the easy version. Ask for the LLP list with remaining cycles and a clean status sheet gives it to you. Ask the same of a photographed tally sheet or a shop visit report scanned at an angle and you get nothing you would put in front of a lessee.

    Context is what turns a fact into a finding. The return conditions annex clause by clause, the delivery condition record from nine years ago, current AD and SB status, the last shop visit report. A cycles figure means nothing until it sits beside the threshold somebody agreed to in 2017.

    Harness is where aviation gets specific. OCR for the scans. A vision model for trend charts, because engine condition data often lives as a plotted line rather than a table. Extraction agents that know an airworthiness review certificate from an LLP back to birth trace from a borescope finding, because each needs different handling. A typed field per return condition, regex verification over the model output, and lineage back to the page number, because you will have to show that page. Then the boundary: the system drafts the discrepancy list, your records lead signs it.

    Loop runs when the back to birth trace has a gap between 2019 and 2021 and one EASA Form 1 is missing. Inferring continuity is the wrong move. Re-search the pack under alternate part numbers, and if it is still missing, raise it as an open item with the serial and the date range attached, so the lessee can be asked for the document while there is still time to get it.

    Memory is knowing that this lessee's MRO issues release certificates under a trade name rather than its legal name, so the matching check fails every time until someone overrides it. Your team worked that out on the last redelivery. If the system needs telling again on this one, you have built a tool rather than a colleague.

    Where the value sits

    Timing is the whole game. Mismatches found on day nineteen arrive after the redelivery date is fixed and the negotiating position has gone, so the lessor absorbs costs that were, on paper, the lessee's. The same mismatches on day three are a conversation while the lessee still has reason to fix things. Nothing goes technically wrong in the day nineteen version. The information existed the whole time, in a format nobody could query.

    We put around forty percent of transition and redelivery costs in the avoidable column, and avoidable versus absorbed is mostly a question of when the gap surfaces. On records audit turnaround, our benchmark is at least sixty percent off, and almost none of that comes from better prompts.

    If your AI project stalled, the model is probably not the reason. Look at what it could see, what it was allowed to do, whether it could notice it was wrong, and whether anything it learned survived until the next run.

    KeepFlying builds the aviation FinTwin on Databricks and Azure, with extraction and knowledge agents tuned to airframe, engine, APU, and landing gear records. Send us a redelivery conditions annex and the records pack that goes with it, and we will show you what we found, what we could not read, and where the system flagged something it was not sure about.

    KeepFlying® builds FinTwin®, the financial twin for aviation — turning airworthiness and maintenance records into validated financial intelligence for Airlines, Lessors and MROs.

    Ready to Transform Your Aviation Operations?

    Discover how our AI-driven solutions can optimize your maintenance operations and reduce costs.