Field notes

What I'd tell a team starting with AI today.

A team beginning with AI faces an appealing distraction: choosing the model feels like choosing the strategy. Model choice matters, especially when privacy, cost or specialist performance is involved. But it is only one decision. The first month is better spent learning where AI can help, what evidence a good result needs and which safeguards let people trust it.

Here is a practical 30-day sequence. It is deliberately modest. The aim is not to transform the company in four weeks. It is to finish the month with one useful workflow, evidence about its limits and a repeatable way to learn.

01

Days 1 to 5

Choose a constraint, then stop shopping

Write down the constraints before comparing products. Which information may the system see? Does data need to remain in a region? What monthly cost is acceptable? Must the model call tools, analyse images or produce code? How quickly must a person be able to review the result? Select a provider, product and deployment arrangement whose documented terms and controls clear those conditions. Record why and keep the choice stable for the pilot unless a security, capability or evaluation failure makes it unsuitable.

This prevents a week of leaderboard tourism. Broadly available model capability is often something you rent. The retained value comes from evaluated examples, approved sources, decisions and the working structure around the model. A new release may later justify a new test. During the pilot, consistency of the experiment matters more than constant switching.

A first AI workflow should earn its next iteration

Compounding begins only when evidence from real use changes the operating system.

  1. Choose one bounded pain Frequent enough to observe, safe enough to test.
  2. Define one measurable outcome Quality, cycle time, error or user effort.
  3. Build the smallest evaluable workflow Prompt, context, tool and check.
  4. Run on representative work Include ordinary and edge cases.
  5. Keep only evidence-backed changes Update the harness or stop the experiment.

A month of measured learning is more valuable than a month of scattered prompting.

Compounding is earnedThe loop compounds when observed failures become better context, controls or evaluations. Mere repetition compounds nothing.
02

Days 6 to 10

Choose one recurring problem worth learning from

Do not start from a demo. Start from work people already do repeatedly and reluctantly. Ask what goes wrong today, who experiences it, how often it happens and what workaround keeps it moving. That is how you walk a requested feature back to the real need.

Suppose an account team spends Friday afternoon preparing renewal briefs. The request might arrive as “build a research agent.” The need is more precise: gather current contract terms, unresolved support issues and recent customer decisions into a reviewable brief, without exposing one customer's information to another. That statement gives the pilot a user, an input boundary, a useful output and a visible risk.

Walk the request back to the pain

A proposed AI feature is useful evidence, but it is not yet the problem definition.

  1. Requested solution For example: build an AI dashboard.
  2. Current workaround Observe what people do today and where effort accumulates.
  3. Consequence Name the delay, error, risk or missed decision.
  4. Underlying need State what must become easier or more reliable.
  5. Smallest responsible response AI is one option, not the predetermined answer.

Start with a consequence that can be observed, then let the solution compete.

Request is not requirementThe quickest way to waste an AI pilot is to automate a proposed interface before understanding the costly behaviour underneath it.
03

Days 11 to 20

Build the smallest workflow you can evaluate

For the renewal brief, begin with five representative, de-identified cases and one reviewer who already knows what good looks like. If real records are necessary, isolate access correctly before the first run. Give the model only approved sources. Define the required sections and require a source beside each material claim. Keep sending, editing and customer contact outside the system. The first workflow should prepare a draft, not make a consequential decision.

Before running it, record how long the current human workflow takes and write a short scorecard: factual accuracy, source freshness, missing risks, cross-customer leakage and minutes of human correction. Define the threshold each measure must clear before seeing the results. Run every example through the same scorecard. When a case fails, record why and change one part of the workflow. This is how a pilot becomes learning rather than a polished demonstration.

A pilot earns expansion through evidence, not applause. Keep the scorecard beside the workflow so checking becomes part of the work.

04

Days 21 to 30

Decide what to keep, change or stop

At the end of the pilot, compare the evidence with the original need. In this hypothetical run, suppose four briefs meet the accuracy and isolation thresholds, while one misses a support risk because the approved source was stale. The sensible decision is not a broad launch. Keep the draft workflow, fix ownership and freshness for that source, then repeat the five-case test. A useful pilot can answer “not yet.” Stopping or narrowing a weak workflow is a successful result if the team keeps the lesson.

If the workflow helps, harden it gradually. Version the instructions and examples. Give each source an owner and review date. Make the evaluation a release gate. Add permissions before adding reach. Finally, write down the trigger that tells the team when to use the workflow. A method left on a shelf is inert; the first skill is recognising the moment it applies.

05

What the first month should produce

A trustworthy lesson is more valuable than a grand launch

Thirty days is not enough to prove an enterprise-wide transformation. It may establish whether one bounded workflow deserves a second month. In regulated or high-stakes work, expert review, an impact assessment and stronger controls can rightly extend the timeline. The useful result is evidence, ownership and a safe next decision, plus a clearer understanding of where the technology bends. When the work genuinely improves, attention can return to the judgement and relationships that needed a person all along.