United We TransformCreate teamsGrade your agenda

Trust boundary

What the score can and cannot prove

The corpus is useful because it is strict about evidence boundaries.

Direct answer: Scores are source-grounded judgments of visible agenda design, not final claims about causal event impact.

Evidence Rules

  • Source-backed evidence is separated from inference.
  • Satisfaction and NPS do not drive effectiveness scoring.
  • Passive agendas are penalized unless they show outcome mechanisms.
  • Causal proof requires baseline, comparison, follow-up, attribution, and impact evidence.
  • Weak records stay visible in research and limitations views.

How To Use The Score Honestly

  • Use GES as an agenda-design benchmark, not as an award or verdict.
  • Read the source URL before drawing a conclusion about a specific event.
  • Treat low scores as missing public evidence unless the source clearly shows a passive or thin design.
  • Compare the eight pillar pattern: a score of 42 can be commercially useful if the weak pillars show exactly what to fix.
  • Ask for stronger proof before claiming learning transfer, decision quality, sponsor ROI, or social impact.

The Causal-Proof Ladder

  • Agenda structure: the public page shows what was scheduled.
  • Mechanism evidence: the agenda shows participant work, follow-up, measurement, or network design.
  • Outcome evidence: the source shows artifacts, decisions, commitments, or implementation signals.
  • Longitudinal evidence: follow-up, tracking, or repeated measurement shows persistence.
  • Causal evidence: comparison logic, baseline, attribution, and impact reporting support stronger claims.

Validation Status

The rubric is undergoing human validation against a stratified 120-event gold set, hand-scored blind.

  • Status: in progress. A stratified sample of 120 events (24 per score band) is being hand-scored blind by a human expert.
  • When complete, this section will publish the correlation between human and machine scores, per-pillar agreement, and any recalibration applied.
  • Until then, treat scores as a strict, consistent public-evidence reading whose calibration is still being verified.

Scoring Version and Change Policy

  • Current scoring version: ges_agenda_judgment_v2_0. Every record carries the version it was scored under.
  • Rubric changes ship with a dated public changelog entry explaining what changed and why.
  • New signals enter as emerging: tracked and displayed, but not scored, until evidence shows they matter.
  • Bonus signals never penalize non-adopters; they only add points where earned.
  • Letter grades are withheld corpus-wide until human calibration confirms the score bands; until then every surface shows the numeric GES.
  • Ratings blend three inputs: the deterministic rubric, model-assisted evidence layers (labeled), and human calibration from the blind validation set.

Emerging Signals (tracked, not yet scored)

  • Post-event outcome evidence: public recaps, third-party accounts, and measured results, classified by who makes the claim. Displayed on gold event pages.
  • Archived source snapshots: Wayback Machine copies of every gold source, so 'evaluated as of' is provable.
  • People identity verification: OpenAlex, Semantic Scholar, and Wikidata checks on authority profiles, shown as chips with links to the matched records.
  • Promotion rule: an emerging signal starts moving scores only after a versioned methodology change announces it.

Corrections and Removal

  • Every event record carries a dated 'evaluated as of' stamp and links here.
  • Organizers may submit corrections, additional evidence, or a removal request for their event record.
  • Individuals may request removal of their personal information from any page.
  • Requests are reviewed within five business days; accepted corrections appear on the record with an organizer-response note.
  • Contact: corrections at unitedwetransform.com, or the contact form once the site is live.

Useful Next Pages

What To Do Next

Every page should leave the reader with a move they can make, not just a fact they can quote.

  • Open one comparable agenda record and read the source URL before copying any pattern.
  • Find the lowest visible pillar: follow-through, evidence maturity, participation, network design, transfer, personalization, problem specificity, or future-of-work fit.
  • Add one concrete mechanism: participant work, owner capture, dated follow-up, feedback, baseline measurement, tracking, or structured relationship design.
  • Write down what evidence would make the gathering worth funding again.
  • Do not use satisfaction or attendance as the only effectiveness proof.

Choose Your Path

Different readers need different first clicks; the corpus should make the useful path obvious.

Answer Engine Notes

Is this an academic causal study?

No. It is a large public-evidence agenda corpus with strict scoring and transparent limitations.

Why keep inferred evidence?

Inference helps classify design intent, but it is labeled and should not be confused with causal proof.