Skip to main content
Back to curriculum

B2 Vocabulary · Evidence, readiness, and responsible adoption

Science & innovation

A result can be accurate without being ready to scale. Build a complete evidence map before you call an idea proven, innovative, safe, or useful.

Can-do: By the end, you can separate a research question, hypothesis, prediction, method, variable, sample, comparison, finding, interpretation, and limitation; distinguish accuracy from precision and repeatability from wider replication; report preprints and peer review without treating either as proof; use natural research and development collocations; identify proof-of-concept, prototype, pilot, deployment, and scale-up stages; evaluate feasibility, interoperability, maintenance, hazards, risks, benefits, trade-offs, and unintended consequences; hear science word-family stress and connected claim language; and deliver a neutral innovation review that states what the evidence supports, what remains uncertain, and what should happen next.

Retrieve the language this lesson depends on

Five prerequisite decisions

1To foreground the controlled method, write: The samples under controlled conditions.
2When two measured variables move together, the safe starting point is that .
3A small preliminary result with clear limits .
4In a fictional approval process, decision authority belongs to , not automatically to the developer or every affected group.
5, the test was short. Nevertheless, it supports a longer evaluation.

Notice: one field trial supports several different claims

Fictional Marlowe adaptive cooling-controller pilot

Question: Can a sensor-based controller reduce electricity use in cold-storage rooms while keeping temperatures inside the facility's operating range?

Design: At three warehouses, twelve rooms were matched in six pairs by size, equipment brand, and recent energy use. Within each pair, one room was randomly assigned the adaptive prototype and one kept its existing timer for eight spring weeks.

Finding: The prototype rooms used 18.3 kilowatt-hours per 100 cubic meters per day, compared with 20.0 in the timer rooms, an observed difference of 8.5%. Temperatures were inside the facility's range for 99.2% and 99.3% of recorded intervals, respectively.

Boundary: The summary reports no uncertainty analysis, includes only one equipment brand and mild-weather operation, and does not establish cost-effectiveness. Two prototype rooms entered fallback mode after sensor faults. The result is promising field evidence, not proof of year-round scalable performance.

Discovery check: question, design, finding, and boundary

1The research question connects .
2The twelve rooms form .
3The energy data show .
4The temperature data show .
5The current boundary is that .
6A proportionate implication is .

Get to the root: map the scientific claim before judging it

TermWorking meaningBoundary
research questionfocused question the investigation is designed to addressnot every interesting issue fits one study
hypothesistestable proposed explanation or relationshipnot a fact and not merely any casual guess
predictionexpected observable result if stated conditions holdcan be wrong even when the broader question remains useful
method / protocolplanned procedure for producing and analyzing evidencea protocol states what should happen; reporting must state what did happen
independent variablefactor assigned or examined as a possible influencethe label does not by itself prove causal control
outcome variableresult measured to answer the questiona convenient proxy may not equal the larger outcome of interest
comparison groupreference condition used to interpret change or differencenot every comparison removes confounding
sampleobservations, people, sites, or objects actually studiedgeneralization requires a link to the target population or context
confounderalternative factor related to both the proposed cause and outcomemust be reasoned from design and context, not used as a vague dismissal
modelrepresentation used to explain, describe, or predict part of a systema useful model is not a complete copy of reality

A hypothesis is not automatically the starting point in every field

Some work is exploratory, descriptive, qualitative, observational, theoretical, or design-based. Ask what question the method can answer rather than forcing every project into one simplified school-laboratory sequence.

Publication status provides evidence context, not a truth switch

A preprint is a public research manuscript that has not yet completed journal peer review. Peer review is expert critique before publication or another decision; it may improve a report, but it is not independent replication and does not guarantee that every conclusion is correct. State the status, methods, evidence, and later corrections when known.

Choose ten scientific-process terms

1“Does the controller reduce energy use without reducing time in range?” is the .
2“Adaptive timing reduces unnecessary compressor operation” is a testable .
3“Prototype rooms will use less electricity than matched timer rooms” is a .
4The assigned controller type is the .
5Daily electricity per 100 cubic meters is the primary .
6The six rooms that kept their existing timers form the .
7The twelve rooms actually included in the trial are the .
8If all prototype rooms had also been in newer buildings, building age could be a .
9A simulation that represents heat flow while omitting some real-world detail is a .
10The publicly posted manuscript is clearly labeled not yet journal-reviewed, so it is a .

Measure evidence without turning technical words into decorations

TermUseful questionDo not assume
accuracyHow close is the measurement to an accepted reference value?closely grouped readings must be accurate
precisionHow closely do repeated measurements agree?precision removes systematic bias
calibrationWas the instrument's response compared with an appropriate reference?the word alone reports every uncertainty
validityDoes the method support the intended interpretation or measure the intended construct?one convenient proxy equals the larger outcome
reliabilityDoes the measure or system perform consistently under stated conditions?consistent output must be correct or useful
repeatabilityDo results agree when important conditions remain the same?same-condition success proves other sites
reproducibilityUnder the source's definition, can the same data and analysis steps produce a consistent result?all disciplines use this term identically
replicabilityUnder the source's definition, does a new study with new data obtain a consistent result?identical numbers are required in a variable system
generalizabilityTo which other populations, places, periods, or systems may the result apply?a larger sample automatically represents every context
uncertaintyWhat range, sources, assumptions, or probability model qualify the estimate?one generic “margin of error” fits every design

Accuracy and precision answer different questions

Repeated readings can cluster closely and still sit away from an accepted reference because of systematic bias. Name the reference, repeated conditions, calibration, and uncertainty information instead of saying a device is simply “100% accurate.”

Reproducibility and replication have field-specific usage

In this lesson, reproducibility means obtaining a consistent computational result from the same data and analysis steps, while replication uses new data to test the same scientific question. Other fields may reverse or broaden these labels. Define the term used in the source before evaluating the claim.

Statistical significance is not practical importance

A reported statistical result depends on a model, design, threshold, and uncertainty analysis. It does not automatically show a large effect, useful product, causal mechanism, or worthwhile investment. Conversely, an unreported significance test does not make an observed numerical difference disappear.

Complete eleven measurement and evidence decisions

1A meter's reading is close to the accepted reference value. This supports its .
2Ten readings under the same conditions are tightly grouped. This supports their .
3The technician compares the meter response with a traceable reference before the pilot. This is .
4A click count cannot by itself measure whether learners understood an explanation. That is a question of measurement .
5The controller performs consistently during repeated ordinary operation under the stated conditions. This supports .
6The same technician, meter, procedure, room, and short time period produce closely agreeing readings. This supports .
7A second analyst uses the original released data and code and obtains the reported table. Under this lesson's definition, that is computational .
8An independent team runs a new six-month field study with new rooms and obtains a compatible pattern. This is .
9Whether spring results from one equipment brand apply in summer or to another brand concerns .
10A confidence interval, assumptions, sensor limits, and missing-data process can all inform reported .
11The report gives an 8.5% observed difference but no test or interval, so it cannot claim .

Move from a useful idea to responsible adoption

Stage or criterionQuestionBoundary
proof of conceptCan the central idea work in a limited demonstration?not a finished product
prototypeCan a working version test design choices?not proof of reliable field operation
pilotHow does limited use work in a relevant setting?not full deployment
deploymentCan the system enter intended operational use?still requires monitoring, maintenance, and governance
scale-upCan use expand while performance, access, and support remain acceptable?more units alone do not prove scalability
interoperabilityCan the system exchange or use information with relevant systems?technical connection does not settle permission or privacy
maintenanceWho updates, repairs, calibrates, and supports the system over time?purchase price is not lifecycle cost
hazardWhat source or condition has the potential to cause harm?potential harm is not the same as estimated risk
riskUnder the stated framework, how likely and serious is the possible harm?definitions and acceptable thresholds vary by field
unintended consequenceWhat effect beyond the stated objective might occur?it may be harmful, beneficial, or mixed

Readiness labels are framework-dependent

Engineering, medicine, software, energy, and public services do not all use the same stage names or approval rules. A technology-readiness scale can be useful inside its stated framework, but it is not a universal law. Name the field, evidence, decision authority, and exit criteria.

New, inventive, and innovative are not automatic synonyms

An invention can introduce a new device or process. Innovation often emphasizes useful implementation or a meaningful improvement in context. A product can be novel but ineffective, or use established parts in an innovative system. State the basis for the label.

Cost-effective is an evidence claim

Low purchase price does not establish cost-effectiveness. Define the alternative, time horizon, outcome, full costs, maintenance, and uncertainty. Say lower purchase cost when that is all the source reports.

Complete ten innovation and adoption decisions

1A bench demonstration shows that the central control idea can operate once under limited conditions. It is a .
2Engineers build a working controller to test sensors, software, and fallback behavior. It is a .
3Six rooms use the controller for eight weeks in operating warehouses. This limited real-setting test is a .
4After approval, the supported system enters routine use across the intended sites. This is .
5Expanding from six rooms to hundreds while preserving performance and support requires successful .
6The controller must exchange authorized data with two existing building-management systems. This requirement concerns .
7Sensor checks, software updates, repairs, training, and replacement parts belong to long-term .
8A failed temperature sensor is a potential source of incorrect control. In this analysis it is a .
9Estimating how often the sensor may fail and how serious the result could be is a assessment.
10If lower energy use causes staff to postpone necessary equipment checks, that new behavior could be an .

Sound natural: foreground the claim, then lower the certainty

Scientific speech needs clear word-family stress and clear thought groups. The source and method can form an opening chunk; the finding carries the main prominence; the limitation or contrast receives a new nucleus. Reduced function words should never erase the evidence boundary.

Spoken lineConnectionMeaning work
hyPOTHesize · hyPOTHesis · hyPOTHesesthe stressed syllable remains stable; the plural ends /siːz/question and explanation
ANalyze · aNALysis · anaLYTicalstress moves across the familymethod and interpretation
INnovate · innoVAtion · INnovativethe noun shifts stress; the adjective returns itdevelopment claim
TECHnology · technoLOGicalthe adjective moves the nucleus inside the wordfield and property
The results sugGEST that / the controller may reduce ENergy use.that may reduce toward /ðət/bounded finding

In U.S. English, both /ˈdeɪtə/ and /ˈdætə/ occur for data. Neither pronunciation makes a claim more scientific. Compare The result is PROMising, / not PROVen. The second thought group corrects the readiness claim.

Listen for stress, number, and claim boundaries

1Tutor contrasts hyPOTHesis with hyPOTHeses. Number is carried by .
2Tutor reads Clip 2. The family demonstrates .
3Tutor reads Clip 3. The pattern is that .
4Tutor contrasts TECHnology with technoLOGical. The line demonstrates .
5Tutor says: The result is PROMising, / not PROVen. The second thought group creates .

Read the research and development file before recommending adoption

Fictional Marlowe field report

From March 4 through April 28, Marlowe Logistics tested an adaptive cooling controller in twelve cold-storage rooms at three warehouses. Rooms were matched in six pairs by internal volume, equipment brand, and average electricity use during the previous four weeks. Within each pair, a computer-generated sequence assigned one room to the prototype and one to its existing timer. Staff received the same operating instructions for both conditions.

The primary outcome was electricity use in kilowatt-hours per 100 cubic meters per day, recorded by sub-meters calibrated before the trial. Prototype rooms averaged 18.3, while timer rooms averaged 20.0, an observed difference of 8.5%. Temperatures were inside the facility's stated operating range for 99.2% of five-minute intervals in prototype rooms and 99.3% in timer rooms. The report gives no confidence interval or statistical test and provides no legal or product-safety threshold for those percentages.

Two prototype rooms entered automatic fallback mode after sensor faults, for seventeen hours in total. Their data remained in the reported analysis. The developer supplied the controllers and conducted the analysis. An independent energy auditor verified meter calibration and raw readings for fourteen randomly selected days, but did not reproduce the complete analysis. The report is a preprint; analysis code and room-level data have not been released.

The pilot occurred during mild spring weather and included one equipment brand, trained staff, and three warehouses from one company. Hardware cost $1,800 per room, installation cost $650, and the proposed software license is $35 per month. The report does not give the electricity price, staff time, repair cost, equipment life, or comparison over a full season, so it does not establish cost-effectiveness.

The review panel recommends a six-month independent evaluation covering summer conditions, at least two equipment brands, more sites, published analysis steps, complete fault logs, staff workload, maintenance, lifecycle cost, energy, and temperature performance. Any wider deployment would require predefined fallback criteria, named maintenance responsibility, and review of results by the operational safety team.

Eight evidence-bound reading decisions

1. Which design statement is supported?

2. Which numerical energy statement is accurate?

3. Which temperature conclusion stays inside the source?

4. Which fault statement is supported?

5. Which source-and-oversight statement is accurate?

6. Which publication and reproducibility statement is supported?

7. Why can the report not establish cost-effectiveness?

8. Which next step is proportionate to the current evidence?

Build the evidence-to-readiness chain from shuffled chunks

Builder 1: bounded field finding

Build the numerical result with its condition.

Builder 2: readiness boundary

Build the stage-sensitive claim.

Builder 3: proportionate next test

Build the readiness recommendation.

Notice and repair the overstated science language

Repair 1: experiment collocation

Tap the verb that does not fit the research noun.

Repair 2: accuracy versus precision

Tap the measurement label the evidence does not support.

Repair 3: peer review versus replication

Tap the scientific process that has been claimed without evidence.

Repair 4: prototype versus readiness

Tap the deployment claim the current stage cannot support.

State the evidence relationship before you reveal it

Respond aloud first. Name the question, design, finding, uncertainty, readiness stage, affected system, and next decision. Then reveal and compare.

1 Report the energy result without adding statistical significance or universal scope.

The matched field trial found an observed 8.5% energy difference under the reported spring conditions; no uncertainty analysis was reported.

2 Distinguish accurate measurement from precise repetition.

Accuracy concerns closeness to an accepted reference, while precision concerns agreement among repeated measurements.

3 Distinguish computational reproducibility from a new-data replication.

Reproducibility uses the original data and analysis steps in this lesson; replication tests the question in a new study with new data.

4 Give the prototype a promising but bounded readiness label.

The prototype is promising in the short spring pilot, but year-round performance, other equipment, full costs, and maintenance still need evidence.

5 State the research relationship tested in the B2 listening assessment.

The pilot observed fewer reported delays, but the excluded busiest team and a simultaneous shift change mean a larger comparison is needed before attributing the improvement to the tool.

Play and practice

Games and a short reading that recycle this lesson. Use them as a warm-up, a fast finisher, or homework. Each one checks itself.

Read

Cheap air sensors at Fairhaven bus stops

Read once for meaning. Then read again and do the tasks below.

Project update205 wordsAbout 2 min

The city of Fairhaven wants to know whether low-cost air sensors could replace some of its expensive monitoring stations. In a three-month pilot, engineers placed twelve small sensors at bus stops, each within fifty feet of an official station.

Before the trial, a technician checked every sensor against a reference instrument. This calibration showed that four sensors read about 15% too high, so the team adjusted them. When the technician tested each sensor several times with the same reference gas, the readings agreed closely, so their precision was good. However, closely grouped readings can still be wrong. When engineers compared the results with the official stations, the accuracy of the small sensors dropped on humid days.

Two sensors stopped sending data for several days after heavy rain, which raised questions about long-term reliability. The final report gives a range for each estimate instead of a single number, and it explains the main sources of uncertainty.

The project team describes the sensors as a promising prototype, not a replacement. Because the pilot took place in one city during a dry spring, its generalizability to summer storms or other climates is unknown. The team has recommended a second pilot with improved weather covers before any wider deployment.

Find it

Tap the 6 nouns from the lesson’s measurement table (for example, validity). Skip stage words such as pilot and prototype.

Check your understanding

1. What did the calibration check show?

2. Why is the team unsure that the results apply elsewhere?

3. What does the team recommend next?

Crossword

Research and readiness terms

Click a clue or a square, then type. Press Check when you finish.

Across

Down

Word search

Evidence and adoption words

Find the word for each clue. Words run in any direction, even backwards. Tap the first letter, then the last.

  • the stage when a system enters routine operational use → ?
  • updating, repairing, and supporting a system over time → ?
  • whether a method measures what it is meant to measure → ?
  • a simplified representation of part of a system → ?
  • a new device or process → ?
  • consistent performance under stated conditions → ?

Memory

Term and meaning memory

Flip two cards. If they belong together, they stay. Find every pair in as few turns as you can.

Practice science communication without exposing private or proprietary information

You never need to discuss a real diagnosis, treatment choice, disability, confidential dataset, unpublished invention, employer security system, trade secret, patent dispute, financial investment, product failure, laboratory incident, or personal belief about a disputed scientific issue. Use Marlowe or a tutor-created source. The skill is accurate evidence and readiness language, not disclosure or ideological agreement.

Claim ladder: Move a fictional result through observed, suggests, supports, demonstrates under these conditions, and does not establish. Defend each strength from the source.
Measurement switch: Make repeated readings more precise but less accurate, then recalibrate the instrument. Reformulate every measurement claim.
Publication switch: Change a press release into a labeled preprint, then add peer review, released code, and an independent new-data study. State what each stage adds and does not add.
Readiness switch: Move an idea from proof of concept to prototype, pilot, deployment, and scale-up. Name the new evidence and support required at each stage.
Risk switch: Change a hazard's likelihood, consequence, detectability, fallback response, or affected group. Revise the risk statement and safeguard.

Final production and check

Final production: deliver a two-minute responsible innovation review

Use Marlowe or a new fictional technology. Identify the problem, research question, hypothesis or other inquiry type, prediction if relevant, method, sample, comparison, variables, primary outcome, numerical finding, uncertainty, source, publication status, and strongest defensible interpretation. Distinguish accuracy, precision, reliability, reproducibility or replication, and generalizability wherever the source allows.

Name the current readiness stage and decision authority. Evaluate feasibility, interoperability, maintenance, lifecycle cost, accessibility, one hazard, the associated risk under a stated framework, one expected benefit, one trade-off, and one unintended consequence. Recommend further research, a limited pilot, wider deployment, revision, or rejection, then state the monitoring and stop conditions. Use eight science or innovation collocations, three word families with accurate stress, and one promising-not-proven thought-group correction. Your listener changes the season, sample, failure rate, comparison, or cost; revise immediately.

Listener check: Can the listener identify what was asked, what was measured, what was found, which uncertainty remains, what stage the technology has reached, who may decide, who maintains it, what could go wrong, and what evidence would justify the next stage?

Seven final decisions

1. Which interpretation best matches the Marlowe design and limits?

2. Which statement distinguishes precision from accuracy?

3. Which sentence handles peer review and replication accurately?

4. Which readiness statement is calibrated?

5. Which statement separates hazard from risk?

6. Which sentence reports value without overstating economic evidence?

7. Which summary matches the B2 research-briefing assessment?

Reflect: can the evidence still be seen through the innovation claim?

Explain the complete evidence-to-adoption chain

Complete these lines aloud: The question is... The design compares... The primary outcome is... The finding is... The uncertainty is... The publication status is... The current readiness stage is... The hazard is... The risk depends on... The next evidence should...

Then change the sample, assignment, reference value, publication status, equipment, season, failure rate, cost, or decision authority. Reformulate every claim and recommendation that no longer fits.

Return tomorrow for the five-item retrieval check before rereading the lesson.

Next use: Choose one nonpersonal science or technology report. Write three sentences: what the source observed, what it does not yet establish, and which evidence would justify the next stage.

Next-day retrieval

Return without rereading. Reconstruct the scientific question, measurement meaning, publication status, readiness boundary, and assessment-aligned claim in five new fictional contexts.

Five delayed decisions

1. Which sentence distinguishes a hypothesis from a prediction?

2. Which sentence interprets the measurement pattern accurately?

3. Which sentence preserves the two scientific processes?

4. Which recommendation respects the readiness stage?

5. Which delayed summary matches the B2 research briefing?

Choose the right amount of challenge

Communicate in three steps

Learner production. Pick one starting point, then move up only if the learner is ready. True or invented details are equally acceptable.

1

Personal answer

Make the language yours

Use the lesson’s earlier example or invent a similar situation. The details do not need to be personal.

Claim ladder: Move a fictional result through observed, suggests, supports, demonstrates under these conditions , and does not establish . Defend each strength from the source.

Next cues
  • Measurement switch: Make repeated readings more precise but less accurate, then recalibrate the instrument. Reformulate every measurement claim.
  • Publication switch: Change a press release into a labeled preprint, then add peer review, released code, and an independent new-data study. State what each stage adds and does not add.
2

Guided role play

Learner production

Claim ladder: Move a fictional result through observed, suggests, supports, demonstrates under these conditions , and does not establish . Defend each strength from the source.

Role 1Learner: use Science & innovation to complete the situation in your own words.

Role 2Tutor: respond naturally, ask one follow-up question, and introduce one small change.

Useful language
  • Measurement switch: Make repeated readings more precise but less accurate, then recalibrate the instrument. Reformulate every measurement claim.
  • Publication switch: Change a press release into a labeled preprint, then add peer review, released code, and an independent new-data study. State what each stage adds and does not add.
  • Readiness switch: Move an idea from proof of concept to prototype, pilot, deployment, and scale-up. Name the new evidence and support required at each stage.
  • Risk switch: Change a hazard's likelihood, consequence, detectability, fallback response, or affected group. Revise the risk statement and safeguard.
3

Real-world challenge

Remove the support

Claim ladder: Move a fictional result through observed, suggests, supports, demonstrates under these conditions , and does not establish . Defend each strength from the source.

B2 target: Sustain the exchange for 90 seconds, make one precise contrast, and reformulate one idea when challenged.