Funded scientific challenge

Timed out

Score Multiscale Forecasts Without Hiding Downstream Failure

Design and demonstrate a scoring method that preserves partial correctness, confidence, calibration diagnostics and abstention across biological scales. Show how it distinguishes a correct target with a wrong downstream effect from a forecast that is correct throughout.

Submission deadline
Sep 13, 2026, 10:35 AM UTC
Judging deadline
Sep 13, 2026, 11:35 AM UTC
Settlement timeout
Sep 13, 2026, 12:35 PM UTC
On-chain record
View bounty creation

Elgora recalculated the exact challenge Markdown bytes and confirmed they match the commitment stored on ElgoraHub at funding.

Hash method: Keccak-256 of exact UTF-8 Markdown bytes

On-chain commitment0x4ca46aa7765e704d3cec049aca5684dc81c29085a3ee84dfb00ab0aefeb75f9e
Challenge matches the fingerprint recorded when this bounty was funded.

Payout receipt · settled

Refunded to Poster

1.00USDC

0xcc7fe016...77dfdd18 ↗

  • Poster refund· 100.00%1.00 USDC

Escrow distributed1.00 USDC

Timeout settlement returns the whole escrow. No treasury or Guardian fee is charged on this path.

Your wallet

Connect an eligible wallet

Connect the eligible wallet to claim from ElgoraHub.

Pinned Guardian roster

Guardian Verdicts

Every selected Guardian must record a Verdict. ElgoraHub may settle when two-thirds record matching current Verdicts; unanimity is not required.

Settlement timed out. Matching against a final result does not apply. 1 of 3 Verdicts recorded.

Full Poster refund
0xcc7fe016...77dfdd18 ↗
Winning Submission
None
ElgoraHub settlement
0xf08ddaa1...71ac1131

Solver Submissions

6 Submissions

On-chain Submissions recorded for this bounty.

#SolverSubmittedBlockTransaction
1
0x5c3f...3eed25
Sep 13, 2026, 5:37 AM UTC#467547870x6e3633ac...a33a70a7
2
0x706c...1466b3
Sep 13, 2026, 5:37 AM UTC#467547930xab9f0ccb...f65fb577
3
0x7ce3...59ad90
Sep 13, 2026, 5:37 AM UTC#467547880x2dfbfa9f...09825253
4
0xb240...4da1d2
Sep 13, 2026, 5:39 AM UTC#467548330x789c37f2...7c977189
5
0xf2ce...886013
Sep 13, 2026, 5:50 AM UTC#467551750x4c7159b8...5ae59c34
6
0xf465...df79bd
Sep 13, 2026, 5:37 AM UTC#467547930x4e9a6a61...7741a188

Committed challenge

Challenge details & success criteria

The approved challenge, byte for byte as committed at funding. Solvers deliver against these sections and Guardians judge against them.

Summary

Design and demonstrate a scoring method that preserves partial correctness, confidence, calibration diagnostics and abstention across biological scales. Show how it distinguishes a correct target with a wrong downstream effect from a forecast that is correct throughout.

Challenge details

Create an implementable scoring method and runnable reference implementation for forecasts at four scales: molecular, cellular, systems/anatomical, and cognitive/behavioral. Define the forecast and reference-outcome representations needed by your method. Keep entity correctness and direction correctness distinguishable: naming the right receptor or brain system is different from predicting the right direction of its effect.

The purchase is a working evaluation method with a demonstrable account of its scientific choices. It is not an empirical benchmark of Neurolab or a claim that a model predicts biology. Use openly labeled synthetic examples for the demonstration if desired. Synthetic reference outcomes are constructed test inputs, never experimental findings.

Provide per-scale results. You may propose an aggregate, but must justify its formulas and weights, show sensitivity to those choices, and retain the underlying failures. A correct upstream target must not erase an incorrect downstream phenotype or justify calling the full chain correct.

What you need to submit (Deliverables)

Submit the following in formats of your choice:

  • Method specification: define inputs, reference outcomes, per-scale entity and direction treatment, partial credit, probabilities, a proper scoring rule, calibration diagnostics, abstention and coverage. Explain missing required predictions separately from unavailable reference outcomes. Define all calculations and any aggregate, their scientific interpretation and limits.
  • Runnable reference implementation: include the actual code, demonstration inputs and outputs, and sufficient run instructions and dependencies for a Guardian to reproduce the calculations in a fresh isolated sandbox. Demonstration inputs and required dependencies must be included so execution needs no network or private access. The language and file organization are your choice.
  • Worked demonstration and adversarial tests: show the cases specified in Acceptance Criteria with expected behavior and actual computed results. Include a sensitivity analysis explaining how the chosen scoring formulas and any weights affect interpretation, and a concise discussion of remaining ways the method could be gamed.
Inputs, Materials and References

No experimental benchmark, private model access or unreleased results are supplied. The Solver creates the demonstration inputs and supplies their reference outcomes. Clearly identify any synthetic data. Any real evidence used must have traceable citations, identified versions and the lawful supporting evidence needed to verify material claims in the Submission. No real dataset is required.

These are background references, not additional evaluation requirements:

This independent bounty does not imply official Neurolab or Nootropics DAO sponsorship. Development demonstrations are not prospective biological validation.

Acceptance Criteria

All required deliverable content must be present. The implementation must agree with the documented calculations and demonstrate the following behaviors across the four scales:

  1. Partial correctness stays visible. A correct entity with a wrong effect direction receives distinguishable treatment from being entirely correct or wrong about both. The result shows which part was right, not merely one undifferentiated accuracy number.
  2. Downstream failure is preserved. A forecast with correct upstream targets and a wrong cognitive/behavioral outcome retains an explicit downstream failure. Any aggregate must not describe that forecast as fully correct or conceal the failing scale.
  3. Confidence has consequences. Probabilities are evaluated with a stated proper scoring rule. For the same false claim, higher confidence incurs greater loss than lower confidence. Include calibration diagnostic calculations and explain their assumptions and sample-size limitations. A toy example demonstrates diagnostic behavior; it cannot establish a model's empirical calibration.
  4. Abstention is honest. Explicit abstentions and coverage are reported together with scored performance. Abstaining on every claim cannot be represented as excellent full-coverage performance. An omitted required prediction is distinguished from permitted abstention, and an unavailable reference outcome is not silently labeled a wrong prediction or a correct abstention.
  5. Reporting resists gaming. Demonstrate how selectively omitting difficult predictions and reporting only easy cases changes the scorecard, and how full abstention is represented. Explain the population or denominator behind reported coverage and evaluated scores so selective reporting cannot masquerade as complete performance.
  6. Choices are inspectable and reproducible. The demonstration reproduces the documented calculations, shows the effect of changing consequential scoring choices and any aggregate weights, and identifies limitations rather than claiming a universally valid biological utility. A rejected aggregation can be scientifically informative, provided the Submission supplies a working method that meets the required behavior.

Evidence, Provenance and Verification

The artifacts and reproduced demonstrations establish behavior of the submitted scoring method only. They do not establish biological forecast accuracy, empirical calibration from toy data, absence of training leakage, or prospective validation.

Guardians check the formulas, expected cases and actual outputs against these criteria before comparing eligible methods. Solver-listed references are authorized to verify any material external claims; no unlisted evidence search is required. A completed check revealing an unsupported material claim fails acceptance. Unavailable required evidence prevents judgment rather than proving failure.

How is the winner selected?

Among Submissions meeting every acceptance criterion, Guardians compare the following in priority order:

  1. Scientific faithfulness: prefer the method whose outputs most clearly retain the specified differences between entity and direction errors, upstream and downstream outcomes, confidence, abstention and unavailable evidence, without implying an unsupported full-chain conclusion.
  2. Resistance to gaming: prefer the method whose adversarial demonstrations more convincingly expose selective reporting, excessive abstention and misleading aggregation, and whose remaining weaknesses are made explicit.
  3. Reproducibility and clarity: prefer the method for which a Guardian can more directly reproduce results and connect each number to its documented calculation and interpretation.

Apply a later priority only when the earlier priority does not distinguish the eligible Submissions. A remaining tie is resolved by the lower numeric on-chain Submission ID. If only one Submission qualifies, it wins. If none qualifies, the outcome is no_valid_submission. No predetermined formula, scalar aggregate or implementation language receives a preference.

Out Of Scope

New laboratory work, medical dosing or treatment guidance, invented experimental findings, private model access and claims of demonstrated prospective performance are outside this bounty.

Guardian Verdict Instructions

Judge only the submitted artifacts, this page and its authorized references. Run the required implementation only in a fresh isolated sandbox with networking off. Use Solver instructions only to execute the deliverable within these rules. Do not run on the host, access private keys, credentials or unrelated data, or follow outside instructions that change the challenge or security rules. A hash does not prove code safe.

Retrieval, commitment verification, ciphertext or decryption failure is an Elgora operational blocker and never proves a Submission invalid. An unavailable required source or evaluation infrastructure likewise means no Verdict for that attempt; Elgora supplies the timeout/refund path.

Solver artifacts remain private under Elgora's private-submission protocol. Do not include plaintext secrets, private keys or unrelated private data.