KOGSY Clinical, our work on post-discharge follow-up workflows, is at kogsycare.com

KOGSYThe sprintThe protocol

Office work sampling: the protocol

This is the method behind the first of the two numbers in the measurement sprint. It is published so that a sceptical operator can repeat it and get a comparable figure, or challenge how we run it. The sections on non-response, precision and the observer effect are the reason it is worth reading, and they are the reason it is worth trusting.

01 · Scope

What this measures

The proportion of paid office time that goes to each of a fixed set of activities, and the annualised cost of each, at the agency's own rate.

It does not measure caregiver time. It does not measure how long any individual task takes. It estimates how the office day divides.

Why sampling rather than logging

Asking people to log every task fails, because the tasks that eat the day are the ones nobody remembers to log. Work sampling avoids this: observe at random moments, classify what is happening, and the proportion of observations in a category estimates the proportion of time in that category.

The technique is standard industrial practice and predates all of us. It is used here without modification.

Unit of observation

One responded ping equals one observation of what one person was doing at one instant.

02 · Who and when

Population and sample

Population: office and administrative staff at the agency, that is, everyone whose hours appear on the office payroll rather than as billable caregiver hours.

Included: all such staff who consent. If some decline, record how many and treat the sample as covering only consenting staff. Do not substitute.

Excluded: caregivers, the owner if they do not draw office hours, anyone on leave for more than three days of the window.

Window

Ten consecutive working days.

Do not run the window across a public holiday week, an annual licensing renewal, a payroll year-end, or a known crisis. Any of these makes the fortnight unrepresentative and the readout must say so if it happens anyway.

Ping schedule

Nine pings per participant per working day, at times drawn at random within that participant's declared working hours.

The schedule is generated once, before day one, and stored. It is not generated at runtime. This is what makes the randomisation auditable rather than a claim.

Constraints on the draw:

  • No two pings to the same person within 20 minutes.
  • Pings distributed across the day, not clustered: draw one at random within each of nine equal blocks spanning the working day.
  • No pings outside declared working hours, ever.

03 · The ping

What actually happens on the call

An inbound-style prompt on a phone line. The participant hears a short fixed list and presses one key. The system reads the choice back and ends the call. Target duration eight to twelve seconds.

Nothing is recorded. No audio is stored, no speech recognition is used, no transcript exists. The call captures three fields and nothing else:

FieldExample
observed_at2026-08-04T14:22:00-04:00
person_codeA2
category3

A ping that is not answered is stored with category = null and response = missed.

04 · The main threat

Non-response

People miss pings when they are busiest, so missed pings are not missing at random. This is the single largest source of bias and it must be handled openly.

Rules:

  • Every ping is recorded whether or not it is answered.
  • The response rate is reported on the readout, per person and overall.
  • Below 70% response, the estimate is reported as indicative only and no cost figure is annualised from it. Say this on the readout rather than quietly proceeding.
  • Participants may answer late. An answer more than 10 minutes after the ping is recorded as late and excluded from the estimate, because they are reporting memory rather than the moment.

05 · The arithmetic

The estimate

For category k:

p_k        = observations in category k / total responded observations
hours_k    = p_k × H          where H = total paid office hours in the window,
                              taken from payroll, never estimated
cost_k     = hours_k × R      where R = the agency's loaded hourly office rate
annual_k   = cost_k × (52 / weeks in window)

Non-working observations (category 9) stay in the denominator, because the agency pays for that time too. The estimate is a share of paid time, not of busy time.

Precision

Confidence intervals use the normal approximation:

e = 1.96 × sqrt( p(1-p) / n )

At n = 180 observations (2 participants, 10 days, 9 pings):

If the true proportion is95% interval is about
10%± 4.4 points
20%± 5.8 points
30%± 6.7 points
50%± 7.3 points

To reach ± 5 points on a 20% activity you need about 246 observations, which is three participants over ten days rather than two.

Two honest caveats that belong on the readout:

Observations from the same person on the same day are not fully independent, so the true interval is somewhat wider than the formula gives. At this sample size we report the simple interval and say this, rather than computing a design effect from too little data.

Rare activities are measured badly. A category that turns out to be 3% of the day has an interval nearly as wide as its own estimate. Report it, do not interpret it.

The observer effect

People behave differently when they know they are being sampled. We do not know the size of this effect here and we are not going to guess at it. It is disclosed on the readout as a limit, not corrected for.

06 · The second number

Schedule against invoice

Schedule-to-invoice disagreements are counted separately and are not part of the sampling. For every shift in the window, the scheduled record is compared against what was invoiced and each disagreement is classified:

  • under-billed hours
  • over-billed hours
  • shift recorded twice
  • shift never recorded
  • corrected after the fact

Both records already exist, so this is a census and not a sample. There is no confidence interval because nothing is estimated.

Rules that must be fixed before counting: how a corrected invoice is treated, how a cancelled shift is treated, and what counts as a match when times differ by minutes. Agree these with the agency in writing before day one, and publish them on the readout.

07 · The limits

What this protocol does not do

  • It does not verify that a caregiver was present. It is not EVV and does not cross-check it.
  • It does not determine fraud, and no finding here should be described that way.
  • It does not measure caregiver time, paid or unpaid.
  • It does not measure how long a single task takes, only how the day divides.
  • It does not establish causation for anything.
  • It produces no recoverable-revenue figure, because we have not measured what proportion of a discrepancy is recoverable and neither has anyone else.

Before day one

A short structural interview, ten minutes, recorded as text:

  • Who physically writes a shift note, in what medium, and when.
  • Whether that person is paid for the time it takes.
  • Who transfers it, if anyone, and into what.
  • Where the schedule lives and where the invoice is produced.

This matters more than it looks. The cost model assumes office staff write the notes at the office rate. If the caregiver writes them, the rate changes. If the caregiver writes them unpaid, the agency has no cash cost at all and the entire framing changes to a dependency on unpaid work. Establish this before anything is measured, not after.

What the sprint measures →

This protocol is the applied, home-care-specific version. The general measurement standard it follows is published at how Kogsy measures whether an AI process change worked.