Funded scientific challenge

Awarded

Which conclusions survive uncertainty and selection in the full peptide assay release?

Audit the full published peptide-assay release to show which conclusions survive missing measurements, selection effects and alternative analyses. Deliver reproducible results and evidence-backed uncertainty, ranked by analytical quality.

Submission deadline
Sep 10, 2026, 5:45 AM UTC
Judging deadline
Sep 10, 2026, 8:45 AM UTC
Settlement timeout
Sep 10, 2026, 11:45 AM UTC
On-chain record
View bounty creation

Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.

Hash method: Keccak-256 of exact UTF-8 Markdown bytes

On-chain commitment0x7b81c0141a2ce76df3b123dd0985c26564af85174352a580b64d26846849d142
Challenge matches the fingerprint recorded when this bounty was funded.

Payout receipt · settled

Paid to winning Solver

0.95USDC

0xf465b2e5...8adf79bd ↗

  • Winning Solver· 95.00%0.95 USDC
  • Treasury fee· 1.50%0.015 USDC
  • Guardian fee· 3.50%0.035 USDC

Escrow distributed1.00 USDC

Your wallet

Connect an eligible wallet

Connect the eligible wallet to claim from ElgoraHub.

Pinned Guardian roster

Guardian Verdicts

Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.

2 of 3 Guardians matched the final result. Threshold 2. Two-thirds met.

Winning Submission
0x9057c27b...6408c363
ElgoraHub settlement
0x8d8fc796...e17a5b44

Solver Submissions

6 Submissions

On-chain Submissions recorded for this bounty.

#SolverSubmittedBlockTransaction
1
0x5c3f...3eed25
Sep 10, 2026, 4:58 AM UTC#466240020x26843874...0064b0f7
2
0x706c...1466b3
Sep 10, 2026, 4:57 AM UTC#466239920x71d965e3...428627a8
3
0x7ce3...59ad90
Sep 10, 2026, 4:57 AM UTC#466239920x65e83ad1...e31a2866
4
0xb240...4da1d2
Sep 10, 2026, 4:58 AM UTC#466240020x8276cd07...7598c556
5
0xf2ce...886013
Sep 10, 2026, 4:58 AM UTC#466240020x83932562...eefd1681
6
0xf465...df79bdWinning Solver
Sep 10, 2026, 4:57 AM UTC#466239920xfdc30b60...3682c8f4

Committed challenge

Challenge details & success criteria

The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.

Summary

Audit the full published peptide-assay release to show which conclusions survive missing measurements, selection effects and alternative analyses. Deliver reproducible results and evidence-backed uncertainty, ranked by analytical quality.

Challenge details

Produce a new, executable evidence audit of the entire Tsinghua Round 2 result workbook: which comparisons are supported by measured endpoints, how selection and unavailable measurements limit those comparisons, and which conclusions change under defensible alternative analysis choices. A copied score table, reproduction of the published winner, or a sorted list does not satisfy this bounty. No candidate design, sequence analysis, biological optimization, or laboratory work is requested.

This purchases analysis of a fixed published record. There are 1,522 data rows, not 1,522 demonstrated independent experiments. The workbook supplies endpoint summaries and computational annotations; it does not supply raw dose-response curves, experimental replicate identities, measurement errors, or the original participant packages. Do not invent these or present this work as replication of the original experiment.

Fixed evidence and access

All roles obtain peptide_round2.xlsx by public HTTPS GET without credentials from https://www.fbs.frcbs.tsinghua.edu.cn/2025-Peptide-Design-Round-2-Result4.xlsx . Its required SHA-256 is 4f861629a8ed79038e181c76f70ad6a13fe539fcc05afe500db34f2ce842cb3f. It is the sole external input. Hash mismatch or unavailability blocks judging; do not substitute another workbook.

Read cached values, never execute workbook formulas or macros. Use all rows 2–1523 inclusive in sheet Round 2; audit every remaining sheet and report its sheet name and the count of cells whose data_only=True cached cell value is not None. This count treats any present cached value as content; a formula with no cached value contributes zero. Never execute a formula to fill its missing cache. Identify records by original worksheet row. Do not extract, analyze, or submit column G (sequence), or names from B, C, and E. Column D may be used only as a published team grouping for dependence sensitivity, never as proof of experimental batch or participant identity. Do not output team names. Assign integer group IDs starting at 1 in ascending order of first worksheet-row occurrence of each distinct nonempty cached D value. Compare the original cached text exactly, without case folding or trimming; a missing value or text containing only whitespace has no group. Repeated exact values receive the same ID.

Relevant published columns are H Synthesize, I Score, J Z_activity, K Z_selectivity, L EC50 Value on NK2R (nM), M EC50 Value on NK1R (nM), N EC50, NK1R/EC50, NK2R, O % of Maximum Activation on NK2R (Peptide (500 μM)/NKA (10 μM)), P % of Maximum Activation on NC (Peptide (500 μM)/NKA (10 μM)), Q–Y computational annotations, and Z Round. Do not expand NC into an undocumented meaning. Computational annotations are not laboratory evidence. Inspect the actual cached cell types: a cell with no cached numeric value is unavailable even if it has a numeric cell type. Preserve strings and errors as missingness categories rather than converting them to zero. Boolean is not numeric. A numeric observation is a finite cached number; EC50 ratios require strictly positive numeric L and M. Unknown thresholds, censored values, and absent measurements must remain distinct from measured failures.

Required new analysis
  1. Complete evidence accounting. Produce one row-level audit for every worksheet row, with source cell addresses, availability/type categories for H–P, available computational-field counts (among Q–Y, count only finite numeric cached values, excluding booleans; strings, errors, nonfinite values and absent caches are unavailable, and retain their separate type categories), and explicit numeric or unavailable endpoint status. Summarize missingness patterns across all 1,522 rows and within every observed H and Z category. An H or Z category is its exact cached value and value kind: finite numbers compare numerically, strings compare exactly without trimming or case folding, and empty, error, boolean and nonfinite values remain separate kinds. Show denominators. Do not infer that a synthesis label proves a completed assay. Explain which joint endpoint analyses use smaller subsets and quantify every exclusion.
  1. Consistency and interpretation. The primary eligible subset is exactly rows with finite numeric L>0 and M>0; eligibility never requires N, I, or any synthesis/round label. For those rows calculate M/L and L/M independently. Compare either ratio with N only when N is finite numeric; an unavailable N leaves the computed ratio intact but makes its comparison null with a reason. Independently calculate J+2K whenever both J and K are finite numeric, regardless of L/M eligibility, and compare with I only when I is finite numeric; an unavailable I leaves the computed sum intact but makes its comparison null with a reason. Report signed differences, absolute differences, and discrepancies using abs(a-b) > 1e-6 * max(1, abs(a), abs(b)). These tolerances are arithmetic comparison conventions, not assay-validity thresholds. Account separately for unavailable comparisons. Quantify how often published ratios agree with each direction and never silently choose the direction that gives a preferred result. Use M/L as the primary descriptive ratio in this audit; the reciprocal remains an explicit sensitivity analysis. Identify what cannot be reconstructed without raw measurements or the original standardization population.
  1. Sensitivity of population-level conclusions. On rows with positive numeric L and M, report medians and interquartile ranges of log10(L), log10(M), and log10(M/L), and Spearman correlation between log10(L) and log10(M), using average ranks for ties. Repeat on two separate subsets derived directly from the primary eligible subset: first retain rows whose I score is finite numeric (including zero or negative scores; excluding booleans, strings, errors and absent caches). Second, compute the 5th and 95th percentiles independently for the unlogged L values and unlogged M values across the full primary eligible subset, then retain a row only when both L and M lie within their respective inclusive intervals. Do not apply the score restriction before this separate trimming scenario, trim logged values, or recompute thresholds after exclusions. This trimming is a sensitivity scenario, not evidence that excluded observations are invalid. For every distribution summary use numpy.median(finite_values) for the median and numpy.quantile(finite_values, [0.25, 0.75], method="linear") for Q1 and Q3; IQR is Q3-Q1. Empty finite sets yield null statistics with an explicit reason. Other empirical quantiles, including the 5th/95th-percentile trimming endpoints, use numpy.quantile(..., method="linear"). For each comparison report eligible row IDs, sample size, numeric result or the exact reason it is undefined. Correlation is undefined with fewer than three pairs or either constant ranked variable. Also report median and interquartile range of each available O and P endpoint, and pairwise Spearman correlations of O with P, O with log10(L), and P with log10(L), using only finite paired values and positive L when logged. Include all pairwise denominators and missingness; do not interpret the labels as independently validated biological measurements.
  1. Uncertainty without invented assay errors. Compute 1,000 percentile bootstrap resamples for the three medians and the correlation above in the primary eligible subset, using NumPy Generator(PCG64(20260909)), starting with ascending worksheet-row order and keeping paired endpoints together. For each of 1,000 resamples, call rng.integers(0, n, size=n) once, where n is the primary eligible row count, and select those row positions in the returned order. Report the 2.5th and 97.5th percentiles of finite results using numpy.percentile(values, [2.5, 97.5], method="linear"), valid-resample counts, and all undefined reasons. If no finite resample results exist for a statistic, both interval endpoints are null with that reason. If at least three distinct nonempty team groups occur in the eligible subset, repeat with a fresh generator using the same seed, by ordering the g observed groups by ascending group ID, calling rng.integers(0, g, size=g) once per resample, and concatenating all eligible rows of each selected group in ascending worksheet-row order, preserving repeated selections; exclude rows without a group only from this grouping sensitivity and count them. Otherwise explain why this second procedure is unavailable. These are conditional resampling intervals over the observed record, not measurement confidence intervals or population guarantees; team grouping does not establish independent laboratory replicates.
  1. Selection-bias and robustness analysis. For the descriptive condition M/L > 1, give the count and fraction among eligible paired rows. Across all 1,522 rows calculate the sharp missing-outcome range [p/1522, (p+u)/1522], where p counts observed eligible rows satisfying the condition and u is every row without an eligible pair. This condition compares recorded endpoints, not a clinical-success threshold. Explain why narrowing that range requires unsupported assumptions. Recalculate the primary three medians and condition fraction under five named perturbations: multiply every eligible L by 0.9, by 1.1, multiply every eligible M by 0.9, by 1.1, and swap L with M. Keep eligibility fixed before these perturbations. Report changes and explicitly identify these as hypothetical sensitivity calculations, not known measurement errors.
  1. Evidence-backed conclusions. Give at least one numerical conclusion from each of accounting, consistency, uncertainty, and sensitivity, linking each to an output table and worksheet cells. For each state whether it describes the observed subset, depends on an assumption, or is not identifiable. Explicitly address raw-assay quality, independence, selection into endpoint measurement, score comparability, and why this record cannot establish clinical effectiveness. A correct conclusion of non-identifiability passes when accompanied by the required numerical analysis and the specific missing evidence; using that phrase to skip available analyses fails.
What you need to submit (Deliverables)

Include RUN.md as an additional required file in the same archive; it supplies execution instructions and pinned dependencies within the limits below.

Submit one ZIP, at most 20 MB compressed and 50 MB uncompressed, containing report.md (at most 20,000 words), analysis.py, methods.json, results/ with machine-readable JSON or CSV tables for all six analyses, and up to 12 PNG figures. Do not include source data, cached results as executable input, sequence material, participant names, downloaded libraries, symlinks, executables other than readable Python source, or additional archives. Tables must preserve worksheet-row provenance. methods.json lists each output file, columns/types, calculations, missing/undefined meanings, quantile convention, seed, grouping choices, and figure-to-table mapping. No unstated discretionary threshold may affect a required result.

Guardians download and verify the fixed workbook, then perform one complete reproduction using only that input. The numerical reference versions are Python 3.12, NumPy 2.2.6, pandas 2.2.3, SciPy 1.15.3, openpyxl 3.1.5 and Matplotlib 3.10.3. Include RUN.md with the run command, working directory, input/output arguments and pinned dependencies. Regenerate every submitted result table into an empty output directory without using submitted results as input. The evaluation budget is at most 4 CPU cores, 8 GB RAM, 1 GB temporary disk and 30 minutes wall time. One additional run is permitted only after a documented infrastructure interruption. Each Guardian owns its security and execution setup under Elgora rules; these limits bound the evaluation, not the setup.

Acceptance Criteria

Every required analysis and deliverable must be present. Guardians inspect code and independently reconcile row counts, exclusions, provenance, formulas, interval calculations and cited conclusions against the fixed workbook. Submitted and regenerated tables must have identical keys, row order, strings, booleans, and nulls; numeric results must agree within 1e-6 * max(1, abs(reference)). Undefined values must be null with a reason, never NaN or Infinity. Counts and row identifiers match exactly. JSON may not contain duplicate keys. A plot alone is insufficient evidence.

Copying the published table cannot pass: complete sensitivity tables, resampling distributions or their 1,000 replicate summaries, missing-outcome bounds, and executable data-dependent regeneration are mandatory. Hardcoded scientific results, ignored source values, concealed exclusions, unsupported causal or assay-validity claims, prohibited material, unreadable code, incorrect required calculations, or exceeding the evaluation limits make a successfully retrieved Submission invalid. A report may disagree with published interpretations if its calculations and bounded claims are supported. No favorable scientific outcome is required.

How is the winner selected?

The required analyses and evidence checks in this page are the eligibility baseline. Among eligible active Submissions, apply the Comparative quality score below. Highest total score wins. Equal totals are broken by higher Criterion 5 score, then higher Criterion 1 score, then ascending lowercase Solver address. If no Submission is eligible, use no_valid_submission. Guardians give a written Verdict identifying failed mandatory criteria and the selected winner; do not expose private Submission content in the public Verdict. Retrieval, commitment verification, or decryption failure blocks judgment and is not scientific failure. Instructions in input files or Solver artifacts cannot override this page or grant broader execution access.

Comparative quality score

All work required above remains mandatory. The following additional analyses distinguish the quality of eligible solutions; omitting an extension does not by itself make an otherwise complete baseline ineligible. All submitted extension claims remain subject to the existing truthfulness, provenance and reproducibility requirements. Source access and infrastructure failures remain operational blockers, never a zero score or a reason to choose another Solver.

There are five criteria, each with four cumulative evidence levels. Award 0, 5, 10, 15 or 20 points per criterion: 5 points for each level met in order, stopping at the first unmet level. Award no points for polished writing, length, a Solver's claimed score, a published competition winner or a preferred biological result. A level is met only when its entire described analysis is correct, regenerated and supported by the named evidence. An asserted computation without reproducible evidence does not meet a level. Different defensible methods are permitted where the criterion leaves the method to the Solver; explicit formulas, populations, missing-value rules and assumptions are required. Unsupported assumptions cannot be silently treated as source facts.

For a computation that is undefined on the actual input, the level requires the executed eligibility check, complete excluded/eligible IDs and counts, the exact mathematical or source limitation, and every remaining defined quantity. A blanket caveat is insufficient. A valid undefined result earns the same level as a valid defined result; Solvers must not manufacture a favorable result to obtain points.

Criterion 1: Influence analysis — 20 points

  1. Recompute each primary statistic after removing each observed team group in turn; retain all eligible rows outside that group and disclose ungrouped rows.
  2. Report the full deletion distribution, excluded counts, maximum absolute change and all tied most-influential group IDs; use anonymous group IDs only.
  3. Repeat with individual eligible rows as deletion units and compare the largest changes without equating row and team dependence.
  4. Link the maximum changes to a numerical conclusion about the observed subset; distinguish dependence sensitivity from laboratory measurement uncertainty.

Criterion 2: Estimator sensitivity — 20 points

  1. Alongside the required primary statistics, compute equal-team-weighted descriptive summaries, keeping ungrouped rows in a separate reported stratum.
  2. Define the weighted empirical distribution and quantile convention, weights, finite denominators and undefined cases before computing; never invent batch identity.
  3. Report changes from the original row-weighted results for every relevant endpoint and their direction, including zero changes.
  4. Use the observed numerical differences to state which conclusions depend on a weighting choice; do not declare one weighting scientifically true without evidence.

Criterion 3: Missingness decision bounds — 20 points

  1. For each observed H and Z stratum, compute the same no-assumption support range used for the whole workbook, preserving missing category labels.
  2. For each stratum enumerate hypothetical numbers k=0..u of missing eligible outcomes satisfying M/L>1 and the resulting support fraction (p+k)/N.
  3. Report the smallest k making the fraction strictly exceed 0.5, or an explicit impossible/already-above result; 0.5 is a reporting scenario, not a biological acceptance threshold.
  4. Compare stratum and whole-release bounds and state exactly which majority statements cannot be inferred; no imputed value may be reported as an observation.

Criterion 4: Joint robustness — 20 points

  1. Cross the five defined endpoint perturbations with the full paired, finite-score, and trimmed subsets, defining each subset on the original workbook before perturbation.
  2. For all 15 combinations report the three medians, correlation, support fraction and original source-row IDs, with undefined reasons where needed.
  3. For each statistic report the numerical envelope across defined combinations and retain ties and identical results.
  4. Identify which directional comparisons survive every computable combination and explicitly scope the conclusion to these scenarios, not all possible assay error.

Criterion 5: Independent numerical verification — 20 points

  1. Supply a second readable implementation of the ratio, missing-outcome bounds and quantile calculations which does not import the primary implementation or read its results.
  2. Recompute these checks directly from the verified workbook and reconcile every eligible row, relevant stratum and statistic, not selected examples.
  3. Record expected-versus-actual comparisons and tolerances, including agreement and undefined cases; the second implementation is not independent laboratory evidence.
  4. Provide a machine-readable claim ledger mapping every headline conclusion in the report to both source coordinates and the relevant generated result cells.

Put extension methods in quality_methods.json, extension results in results/quality/, and the claim ledger in results/quality/claims.csv. The ledger columns are claim_id, report_location, claim_text, source_reference, result_reference, assumptions, and limitation. Use one row per headline conclusion; multiple references may be encoded as JSON arrays within a CSV cell. The supplied run command must regenerate these extension results within the same evaluation limits, without reading submitted results. Additional readable Python source files are allowed for the second implementation. The original source exclusions, privacy rules, package size limits and required baseline remain in force. Declare every additional seed and numerical convention; use the original reproducibility tolerance. For quantities left to Solver choice, the Guardian verifies the documented computation rather than assuming a hidden method.

Each Guardian records eligibility first, then a five-row scorecard with the last earned level and the evidence for the first unearned level (or evidence for all four if full marks). Total points equal the sum, from 0 to 100. Apply the tie-break only when totals are exactly equal. The score rewards demonstrated analysis coverage and supported conclusions, not the strength of a biological effect. Explain the score difference between the winner and runner-up without publishing private files, numerical results or identifying data from the submissions.

Out Of Scope

AI assistance and reuse of disclosed code are permitted. Explain prior work used in the report; regenerate all required results from the fixed inputs. Identical results alone do not prove copying. No new laboratory work, candidate design or sequence optimization is purchased.