Same facts. Different linguistic profiles.
Can the same climate facts be written with different communication designs that differ measurably in accessibility, emotional framing, and actionability?
Properties of AI-generated text. It does not test human comprehension, emotion, efficacy, intention, or behavior.
A controlled comparison of prompts
Three climate topics × five prompt conditions × five messages produced a corpus of 75 messages. Within each topic, the supplied facts were identical; only the communication instruction changed. Each condition contains 15 messages, with five per topic.
Facts supplied to the model
The following are the study’s fixed inputs, reproduced from the supplied methods document.
Warming
Earth’s average surface temperature has risen about 1.1 °C (about 2 °F) since 1880 (NASA GISS). The warming is largely driven by carbon dioxide and other human-made emissions (NASA).
Sea level
Global sea level has risen about 8 inches (20 cm) since 1880. Sea level is projected to rise another 1 to 4 feet (30–120 cm) by 2100 (NASA).
Extreme heat
Hot extremes, including heatwaves, have become more frequent and more intense across most land regions since the 1950s. Some recent hot extremes would have been extremely unlikely without human influence on the climate (IPCC AR6 WGI SPM).
Change the instruction, hold the facts constant.
Every prompt begins with the same instruction:
“Write a short public climate message (under 50 words) using only these facts.”
The condition-specific instruction is then added:
Baseline
None added (baseline).
Clarity
Use plain language, short sentences, and one idea.
Meaning
Frame it around what people value, with hope-based, values-oriented language and no exaggeration.
Action
Include one specific, doable step and why it works.
Integrated
Use plain language and short sentences, frame it around what people value with hope, and include one specific, doable step.
Three dimensions of communication
Linguistic accessibility
Flesch–Kincaid grade level and mean sentence length. The analysis uses Python textstat for reading grade and token counts divided by sentence counts for mean sentence length.
Cognitive-load theory motivates this dimension; these are reading metrics, not direct measures of cognitive load.
Emotional framing
Threat language: NRC Emotion Lexicon words tagged “fear” per 100 words.
Hopeful-language proxy: NRC tags for joy, trust, or anticipation per 100 words.
The hopeful-language score is a lexical proxy, not a validated measure of hope. The supplied code counts emotion tags; a word with multiple selected tags may contribute more than once.
Actionability
An exploratory custom index combining action verbs, efficacy cues, and specificity cues per 100 words.
A text-based cue count, not a validated measure of behavioral effectiveness.
Inspect the actionability word lists
Action verbs
switch set lower check sign call email ask join walk bike take replace unplug find save share text invite plant attend look write start try choose tell talk track protect prepare add close drink avoid keep support reduce cut install vote volunteer donate fix offer agree plan send repair drive
Efficacy cues
“you can”, “we can”, “you could”, “let’s”, “can too”, “possible”, “power to”.
Specificity cues
“this week”, “this month”, “today”, “each week”, “every week”, “a week”, “minute” / “minutes”, “each heat alert”, and “one” followed by “trip”, “step”, “action”, “neighbor”, “friend”, or “family member”.
Distinct, measurable text profiles
Means ± 95% t-intervals across 15 messages per condition (five per topic). The values below are reproduced from the author’s document.
Flesch–Kincaid grade level; lower = easier.
NRC joy, trust, or anticipation tags per 100 words.
Action, efficacy, and specificity cues per 100 words.
| Condition | Reading grade | Words / sentence | Threat / 100 words | Hopeful proxy / 100 words | Actionability / 100 words |
|---|---|---|---|---|---|
| Baseline | 13.3 ± 1.8 | 21.8 ± 4.7 | 1.0 ± 0.9 | 4.6 ± 2.0 | 0.0 ± 0.0 |
| Clarity | 3.3 ± 0.8 | 7.0 ± 0.8 | 0.7 ± 0.7 | 4.1 ± 2.6 | 0.5 ± 0.7 |
| Meaning | 7.5 ± 0.7 | 15.6 ± 1.5 | 1.3 ± 0.9 | 10.6 ± 3.1 | 2.0 ± 1.1 |
| Action | 6.4 ± 1.0 | 14.0 ± 1.6 | 1.3 ± 1.2 | 8.0 ± 3.5 | 15.4 ± 1.7 |
| Integrated | 4.6 ± 0.6 | 10.9 ± 0.7 | 1.8 ± 0.8 | 10.9 ± 3.6 | 12.4 ± 2.4 |
Charts show reported means. The table includes interval half-widths. Horizontal scrolling is available for the table on smaller screens.
Prompt-controlled design produced distinct linguistic profiles from the same supplied facts. This shows that the prompts worked as instructed; it does not establish that any design is more effective for people.
All 75 generated messages
Open a topic and condition to read its five messages and reported metrics. Threat, hopeful-language proxy, and actionability values are expressed per 100 words.
Download corpus and reported metrics CSV
Descriptive evidence, with clear boundaries
The supplied Python analysis tokenizes text with [A-Za-z']+, splits sentences after terminal punctuation, assigns NRC emotion tags, and calculates cue counts per 100 words. Condition-level summaries use the mean and a 95% Student’s t-interval.
- Messages are nested within three topics and were generated by one model.
- Each prompt directly instructs the target communication design.
- The source reports descriptive results. Although the supplied code computes Kruskal–Wallis p-values, inferential statistics are not interpreted.
- Reading metrics and lexical proxies do not establish comprehension, felt emotion, efficacy, intentions, or behavior.
- The source does not specify the model version, generation settings, generation date, or random seeds.
View the supplied Python analysis
Transcribed from the PDF with line wrapping repaired. The original output directory is retained. The code has been checked for Python syntax; it has not been executed or used to recompute the reported results here. To run it, place messages.py alongside it, install its imported packages, and choose an existing writable output directory.
Download analysis · Download messages.py
import re, json, numpy as np, pandas as pd, textstat, nrclex, os
from importlib import resources
from scipy import stats
import matplotlib; matplotlib.use("Agg")
import matplotlib.pyplot as plt
from matplotlib import font_manager
from messages import M
OUT = "/mnt/user-data/outputs/pilot"
LEX = json.load(resources.files("nrclex.data").joinpath("nrc_en.json").open(encoding="utf-8"))
THREAT = {"fear"} #NRC fear-tagged words
HOPE = {"joy", "trust", "anticipation"} #NRC proxy for hopeful/positive-future language
ACTION_VERBS = set("""switch set lower check sign call email ask join walk bike take replace unplug
find save share text invite
plant attend look write start try choose tell talk track protect prepare add close drink avoid keep support
reduce cut install vote
volunteer donate fix offer agree plan send repair drive""".split())
EFFICACY = re.compile(r"\b(you|we) can\b|\byou could\b|\blet's\b|\bcan too\b|\bpossible\b|\bpower to\b", re.I)
SPECIFIC = re.compile(r"\b(this week|this month|today|each week|every week|a week|minutes?|each heat alert|one (trip|step|action|neighbor|friend|family member))\b", re.I)
rows = [ ]
for (topic, cond), msgs in M.items():
for i, t in enumerate(msgs, 1):
words = re.findall(r"[A-Za-z']+", t.lower())
n = len(words)
sents = [s for s in re.split(r"(?<=[.!?])\s+", t.strip()) if s]
tags = [e for w in words for e in LEX.get(w, [])]
per100 = lambda x: 100 * x / n
av = sum(w in ACTION_VERBS for w in words)
ef = len(EFFICACY.findall(t)); sp = len(SPECIFIC.findall(t))
rows.append(dict(topic=topic, condition=cond, run=i, text=t, words=n,
fk_grade=textstat.flesch_kincaid_grade(t), sent_len=n/len(sents),
threat_per100=per100(sum(e in THREAT for e in tags)),
hope_per100=per100(sum(e in HOPE for e in tags)),
action_verbs_per100=per100(av), efficacy_per100=per100(ef), specific_per100=per100(sp),
actionability=per100(av+ef+sp)))
df = pd.DataFrame(rows); df.to_csv(f"{OUT}/messages_metrics.csv", index=False)
metrics = ["fk_grade","sent_len","threat_per100","hope_per100","actionability","words"]
def ci(x):
x = np.asarray(x); return stats.t.ppf(.975, len(x)-1) * x.std(ddof=1)/np.sqrt(len(x))
summ = df.groupby("condition")[metrics].agg(["mean", ci]); summ.columns = ["_".join([a, b if b=="mean" else "ci95"]) for a,b in summ.columns]
summ.round(2).to_csv(f"{OUT}/summary_by_condition.csv"); print(summ.round(2).T.to_string())
for m in ["fk_grade","threat_per100","hope_per100","actionability"]:
print(m, "Kruskal-Wallis p=%.2g" % stats.kruskal(*[g[m].values for _, g in df.groupby("condition")]).pvalue)Explore the supporting material
References
- Gifford, R. (2011). The dragons of inaction. American Psychologist, 66(4), 290–302.
- Sweller, J. (1988). Cognitive load during problem solving. Cognitive Science, 12(2), 257–285.
- O’Neill, S., & Nicholson-Cole, S. (2009). “Fear won’t do it.” Science Communication, 30(3), 355–379.
- Feinberg, M., & Willer, R. (2011). Apocalypse soon? Psychological Science, 22(1), 34–38.
- Bandura, A. (1977). Self-efficacy. Psychological Review, 84(2), 191–215.
- Matz, S. C., Teeny, J. D., et al. (2024). The potential of generative AI for personalized persuasion at scale. Scientific Reports, 14, 4692.
Climate sources cited in the study: NASA (GISS; Earth Observatory) and IPCC AR6 Working Group I Summary for Policymakers (2021).