Analysis · Nebula Components

A Green Lighthouse Score Is Not a Field Performance Baseline

How to use lab diagnostics and field evidence together when judging landing-page performance.

AnalysisMeasurementEvidenceField performance
A Green Lighthouse Score Is Not a Field Performance Baseline

A strong Lighthouse result is useful. It is not the same thing as knowing how real visitors experience a page.

That distinction matters when a founder is deciding whether a landing page is ready for a campaign. If the lab score is the only evidence in the room, the team can make a confident decision from a narrow observation. The practical answer is to use lab data to investigate and field data to judge the experience.

Lab and field performance answer different questions A controlled lab test helps isolate causes. Field evidence shows how real visitors experience a page across devices and networks. A reliable decision uses both. LAB FIELD Repeatable setup Isolates likely causes Diagnostic signal Real devices and networks Shows visitor experience Decision evidence Use lab evidence to investigate, then return to field evidence.
Lab data helps explain what to investigate. Field data tells you what visitors experienced.

The two measurements answer different questions

Lab data is collected under a controlled setup. The device, network, browser conditions, URL, and run are defined so that a change can be compared with another run. That makes lab testing useful for diagnosis. You can ask whether a large image, a blocking script, or a layout change creates a repeatable problem.

Field data is collected from real visits. Devices, connections, page states, and user behavior vary. The result is a view of what visitors experienced across a population and a time window, not a single controlled run.

Neither measurement is a more complete replacement for the other. They describe different parts of the problem. A lab result can be repeatable without being representative of every visitor. A field result can be representative without telling you which line of code caused the experience.

The distinction is easy to lose because both outputs often arrive as scores, percentiles, or pass and fail labels. A score compresses the underlying observations. Before making a decision, recover the scope behind the score: what was measured, under what conditions, for which visitors, and during which period?

Why a green lab result can still leave a campaign question unanswered

A controlled run can tell you that a page behaved acceptably under the selected conditions. It cannot, by itself, tell you that every important visitor segment received that experience. A campaign may bring traffic from older phones, slower connections, different geographic regions, or a different landing-page state than the test used.

This does not make lab testing unhelpful. It makes the test a diagnostic instrument rather than a universal certificate. A green result can narrow the investigation. It may tell you that a suspected change did not reproduce a problem in the test setup. It may also reveal that the test needs to be repeated with a different page state or a more representative configuration.

The same caution applies to a poor lab result. A poor controlled result is a reason to investigate a plausible technical cause. It is not automatically proof that a campaign will fail, nor is it proof that changing the page will improve a business outcome. The business decision still needs a link to the relevant field population and observation window.

How to read disagreement

When lab and field results disagree, do not start by choosing the result you prefer. Classify the disagreement first.

Lab resultField resultUseful interpretation
PoorPoorThere is a strong reason to investigate. Use the lab to isolate likely causes, then use field data to understand who is affected.
PoorHealthyThe controlled setup may expose a diagnostic risk that is not material for the observed population, or the field window may not capture it. Check the scope before changing the page.
HealthyPoorThe lab setup may be too narrow or too different from real traffic. Investigate device, connection, page-state, and population differences before trusting the green score.
HealthyHealthyThe evidence is reassuring for the measured scopes and time windows. It is not proof that every future visitor will have the same experience.

This table is not a scoring system. It is a way to stop a single number from carrying more certainty than the evidence supports. The next question is always whether the two measurements cover the same thing.

Compare the URL and template first. Then compare the page state, device mix, network conditions, geography, traffic source, and dates. A disagreement caused by different scopes is not necessarily a contradiction. It may be two accurate observations of different situations.

A review sequence that leads to a decision

1. Start with the field question

Write down what you need to know about real visitors. For example: are mobile visitors experiencing a meaningful delay on the campaign landing page, and during what period? Define the population and time window before interpreting the result.

If no field baseline exists, record that as an information gap. Do not describe a lab score as a substitute for evidence about visitors. A missing baseline is a decision constraint, not a reason to invent one.

2. Use lab testing to investigate

Run the lab test with a known setup and record the URL, device and network conditions, date, and relevant result. Look for a plausible cause rather than treating the score as a verdict. A repeated result can help identify an oversized image, delayed script, unstable layout, or another concrete investigation target.

Change one meaningful factor at a time when possible. If several variables change together, a better score will not tell you which change helped. Keep the before and after conditions close enough that the comparison remains useful, and record anything that changed outside the intended intervention.

3. Return to field data

After the change, check the field signal again over an appropriate window. Ask whether the population or device segment that mattered actually improved. If the field signal is unavailable or too small to support a conclusion, say so plainly.

The useful conclusion may be narrower than “the page is fast now.” It might be: “the controlled run improved after the image change, but we do not yet have enough field evidence to determine the effect on visitors.” That is a better decision boundary because it keeps the known result separate from the unknown one.

A worked example without invented numbers

Imagine a paid campaign is about to send mobile traffic to a landing page. The team runs a controlled test and sees a poor result associated with the page's largest image. The test gives the team a concrete investigation target: inspect the image dimensions, format, loading behavior, and whether it is needed above the fold.

The team makes one image change and repeats the same controlled test. The result improves. That is useful evidence that the controlled condition changed. It is not yet evidence of how campaign visitors responded.

The team then checks the field signal for the campaign's mobile population during a defined window. Three outcomes are possible. If the field signal improves in the relevant segment, the intervention has support from both diagnostic and visitor evidence. If the field signal does not improve, the image may not have been the material cause for those visitors, or another bottleneck may dominate. If the field signal is unavailable, the correct conclusion is that the intervention is awaiting field confirmation.

Notice what this example does not claim. It does not claim a percentage improvement, a conversion lift, or a universal performance result. Those claims would require actual measured evidence. The value of the process is that it tells the team what to inspect next and what conclusion the current evidence can support.

What to record so the comparison stays honest

For each review, keep a small evidence note with:

  • the exact URL and page or template version
  • the lab device, network, browser, and run date
  • the field population and observation window
  • the result or signal being compared
  • the change made between runs
  • what the evidence supports
  • what remains unknown

This context matters because a result without its scope is easy to reuse as if it were universal. A field observation from one segment does not automatically describe every visitor. A single lab run does not establish a durable baseline.

Also record whether the comparison is intended to answer a diagnostic question or a visitor-experience question. That one sentence prevents later readers from treating a diagnostic result as a business outcome.

A practical campaign-review worksheet

Before approving a campaign landing page, answer these questions in order:

  1. What visitor population matters for this campaign?
  2. What time window will make the field observation meaningful?
  3. What field signal will indicate that the experience needs investigation?
  4. What controlled test can help isolate a likely cause?
  5. Which single intervention will be evaluated first?
  6. What result would count as improvement in the controlled setup?
  7. When will the field signal be checked again?
  8. What conclusion is allowed if the field evidence is missing or inconclusive?

The worksheet is deliberately modest. It does not turn performance review into a ceremony. It makes the evidence boundary visible before a green score gets reused as a campaign guarantee.

What different outcomes mean for the next action

If the lab and field results are both poor, prioritize an investigation that can produce a concrete change, then schedule a field recheck. Do not stop at the score. Name the suspected cause and the observation that will tell you whether it mattered.

If the lab is poor but field evidence is healthy for the relevant population, check whether the lab setup is materially different before spending effort on a broad page change. A lab issue may still matter for another audience, but the current campaign decision should use the scope that matches its visitors.

If the lab is healthy but field evidence is poor, treat the green result as insufficient. Reproduce the field conditions as closely as possible, inspect the page state reached by visitors, and look for differences in device, network, geography, or traffic path.

If both are healthy, document the scope of that reassurance. Keep monitoring the relevant population when the campaign, template, assets, or scripts change. Healthy evidence supports a bounded conclusion. It does not create a permanent guarantee.

How to communicate the result to a team

A useful performance note can fit into four parts:

  • Observation: what each measurement actually showed.
  • Scope: which visitors, setup, URL, and dates were included.
  • Interpretation: what the evidence supports and what it does not support.
  • Next action: the smallest investigation, intervention, or field recheck warranted.

For example, a note might say that the controlled run identified a repeatable image-loading issue, the image was changed, and the controlled result improved. It would then state that field confirmation for the campaign population is still pending. That language is more useful than saying the page is now fast, because it gives the next person a precise place to continue.

What this means before a campaign

A green lab score is a useful reason to look closer, not a reason to stop looking. Before sending more paid traffic to a page, use the evidence you have to answer three separate questions:

  1. Is there a real visitor problem in the population that matters to this campaign?
  2. Does a controlled test provide a credible explanation worth investigating?
  3. After the change, what evidence shows that the relevant visitor experience improved?

If the answer to the first question is unknown, do not manufacture certainty from the lab score. If the answer to the second is yes, use the lab result to choose a focused investigation. If the answer to the third is not yet available, label the change as an intervention awaiting field confirmation.

The lab explains what to investigate. The field tells you what visitors experienced. A defensible performance decision needs both scopes kept distinct. That is the practical lesson: use the controlled result to make the next investigation smaller, and use the field result to decide whether the change mattered to the people you actually need to serve.

Sources

Why lab and field data can be different

Core Web Vitals workflows with Google tools