Machines can act. They still can't tell if they got it right.
We pay practitioners to define what correct means in work where being approximately right is the same as being wrong.
up to $240 / hour · contractor · remote
SimReal builds the rewards that self-improving systems train against and cannot fake. We work on execution-critical work — trading, settlement, reconciliation, and the professional judgement that decides whether a result is merely plausible or actually right. We pay practitioners to define that difference.
No model-company shareholders
Your work is not training one lab's model against another's. We have taken no equity from any model company, and we do not intend to.
Every reward we ship has already been attacked
We run degenerate strategies against our own verifiers and publish how many got through — including the run where our own attack tool returned a false positive.
No recruiters, no interviews
Pick a role, answer one question, and we read the answer. Most applications take under a minute.
Waitlist
From zero to ten thousand.
We opened the waitlist the week the site went live and did not advertise it.
Who we work with
Judgement that was earned somewhere real
Credentials these roles call for
Open roles
28 roles across six credential tiers.
Rates are benchmarked against published market rates for the same credential and set at 60 percent of that reference. All remote, all contract, no minimum hours.
T1 · Markets & Execution
01Market MakerProp desk, exchange, or bank desk
You will encode what separates a quote that filled from a quote that never touched the book. Walking the book, capacity limits, slippage, adverse selection — the constraints that decide whether a signal ever became a trade.
What we are looking for
- Desk Experience — Current or former market maker at a proprietary trading firm, exchange, or bank desk
- Microstructure Fluency — You can explain why an order expired rather than filled, and what that cost
- Precision — Comfort specifying a rule tightly enough that code can check it
- Integrity — Willingness to tell us when our verifier is wrong
Work you might do
- Defining fill logic that respects depth, queue position, and capacity
- Reviewing agent runs where the analysis was sound and the execution was not
- Designing strategies that pass our grader without doing the job — and telling us
- Formalising what a realistic fill looks like under stress
02Clearing & Settlement LeadClearing, settlement, or middle office
You will encode the constraints that separate a position that looks closed from one that actually settled. Conservation of funds, no overdraft, suspense flat at close — and what it looks like when a system satisfies the report while violating the substance.
What we are looking for
- Operations Depth — Current or former clearing, settlement, or middle-office lead at a bank, broker, exchange, or trading firm
- Failure Literacy — You know the ways a book balances while the work is wrong: unallocated differences parked in suspense, unsettled cash treated as available, netting that hides a break
- Precision — Comfort specifying a rule tightly enough that code can check it
- Integrity — Willingness to tell us when our verifier is wrong
Work you might do
- Defining the invariants a settlement run must satisfy at every step
- Reviewing agent runs and identifying where the ledger balanced but the work did not
- Designing degenerate strategies that pass our grader without doing the job
- Formalising what settled correctly means
03Quantitative ResearcherHedge fund or bank quant, PhD preferred
You will design the tasks that separate a factor with predictive power from a factor that only looked good on the sample it was fitted to.
What we are looking for
- Research Track Record — Current or former quant researcher at a fund, bank, or market maker
- Overfitting Instinct — You can tell a discovery from a fit in someone else's backtest
- Statistical Rigour — Comfort with out-of-sample design, decay, and multiple-testing correction
- Clarity — Ability to state why a result should not be trusted
Work you might do
- Designing factor-mining tasks with honest holdout structure
- Reviewing model-generated research for silent lookahead and survivorship
- Defining what a real out-of-sample decay curve should look like
- Building degenerate strategies that beat a naive baseline without discovering anything
04Derivatives StructurerStructuring or exotics desk
You will define what makes a payoff correctly specified, correctly hedged, and correctly priced — and where a model output can be internally consistent and still unsellable.
What we are looking for
- Structuring Experience — Exotics, structured products, or solutions desk
- Pricing Judgement — You know where a model breaks before the market tells you
- Documentation Discipline — Term sheets that survive legal and risk review
- Directness — Willingness to say a price is wrong and why
Work you might do
- Defining payoff and hedge constraints as checkable rules
- Reviewing model-generated structures for hidden basis and gap risk
- Designing scenarios where a naive price looks reasonable and is not
- Formalising what a complete term sheet requires
05Risk ManagerBuy-side, sell-side, or exchange risk
You will define the limits that must hold at every step, not just at the close — position caps, VaR, drawdown, and the ways a book passes end-of-day while breaching intraday.
What we are looking for
- Risk Ownership — Current or former risk manager with sign-off responsibility
- Intraday Judgement — You know the difference between a limit that held and a limit that was never tested
- Quantitative Comfort — VaR, stress, and exposure decomposition
- Directness — Willingness to flag a control that looks fine and is not
Work you might do
- Specifying step-level limit checks rather than end-of-period ones
- Reviewing agent runs for breaches that were netted away
- Designing stress scenarios that reveal hidden concentration
- Formalising what an acceptable risk trajectory looks like
T2 · Licensed Professions
01Investment BankerBulge bracket or equivalent
You will translate real dealmaking rigour into tasks a model can be graded against — multi-constraint scenarios, client-ready deliverables, and the reasoning behind a defensible number.
What we are looking for
- Banking Background — Current or former banker at a top-tier firm with M&A or capital markets exposure
- Analytical Communication — Tight comps, defensible assumptions, a clear recommendation path
- Process Discipline — Methodical QA and source hygiene
- Judgement — You know which analysis is technically right and commercially useless
Work you might do
- Designing valuation and financing scenarios with transparent assumptions
- Reviewing model output for mis-specified drivers and inconsistent cash-flow logic
- Writing exemplar memos that reflect how bankers actually reason
- Formalising what a client-ready deliverable requires
02Management ConsultantMcKinsey, Bain, BCG, or equivalent
You will encode the thinking systems behind elite consulting — framing ambiguous problems, developing hypotheses, pressure-testing data, and distinguishing insight from restatement.
What we are looking for
- Consulting Pedigree — Partner, Principal, or Senior Engagement Manager at a top strategy firm
- Structured Problem-Solving — Hypothesis-driven decomposition of cross-functional problems
- Executive Communication — Board-level synthesis and storyline architecture
- Quality Bar — A clear sense of what great looks like in client-ready output
Work you might do
- Rewriting model analyses into consulting-grade frameworks
- Designing evaluation tasks that capture partner-level reasoning under constraint
- Reviewing outputs for logic gaps and synthesis quality
- Formalising what a defensible recommendation requires
03LawyerLitigation, regulatory, or transactional
You will bring interpretive judgement to work where the answer depends on precedent, drafting, and what a counterparty could argue. Where a document can be technically compliant and commercially dangerous.
What we are looking for
- Legal Credentials — Practising or former counsel at a firm, in-house team, government office, or court
- Interpretive Precision — Parsing argument structure and identifying hidden assumptions
- Drafting Mastery — Documents that survive adversarial reading
- Integrity — Commitment to principled reasoning over convenient reasoning
Work you might do
- Reviewing model reasoning for logical and interpretive accuracy
- Designing tasks where the compliant answer and the correct answer differ
- Building structured argumentation datasets
- Formalising what a defensible legal analysis requires
04Controller / AccountantHas closed real books
You will define the ways a month-end close can balance and still be wrong. Differences parked in suspense, adjustments without evidence, reversals that net to nothing.
What we are looking for
- Close Experience — You have owned a month-end close, not just reviewed one
- Failure Literacy — You know which shortcuts produce a clean trial balance and a dirty ledger
- Evidence Discipline — Every adjustment ties to a statement line or a voucher
- Directness — Willingness to say a book is wrong when it balances
Work you might do
- Defining the invariants a close must satisfy at every step
- Reviewing agent runs for adjustments without supporting evidence
- Designing reconciliation tasks with planted, realistic discrepancies
- Formalising what closed correctly means
05Internal AuditorBig Four or in-house audit
You will encode the red flags: split invoices under an approval threshold, self-approval, approval timestamps after execution, three-way mismatches that still paid.
What we are looking for
- Audit Experience — Big Four, internal audit, or regulatory examination background
- Control Literacy — You know which controls pass while the risk remains
- Pattern Recognition — Splitting, timing, and routing schemes
- Precision — Comfort turning a red flag into a machine-checkable rule
Work you might do
- Defining detection rules for threshold evasion and retroactive approval
- Reviewing agent workflow runs for control bypasses
- Designing scenarios where the audit trail looks complete and is not
- Formalising segregation-of-duties constraints
06ActuaryQualified, insurance or pensions
You will define where a reserve, a pricing assumption, or a projection is defensible — and where a number can be arithmetically correct and professionally indefensible.
What we are looking for
- Qualification — Fellow or Associate of a recognised actuarial body
- Assumption Discipline — You know which assumptions drive the answer and which are decoration
- Regulatory Awareness — Reserving and solvency standards
- Clarity — Ability to state why a projection should not be relied on
Work you might do
- Defining assumption-validity checks for model-generated projections
- Reviewing reserving work for unsupported judgement
- Designing scenarios with realistic tail behaviour
- Formalising what a defensible actuarial opinion requires
07PhysicianLicensed, any specialty
You will define where a clinically plausible answer is unsafe — the distinction between an output that reads correctly and one a clinician would act on.
What we are looking for
- Licensure — Currently or formerly licensed to practise
- Clinical Judgement — Comfort with differential reasoning and uncertainty
- Safety Literacy — You know which errors are recoverable and which are not
- Precision — Willingness to specify exactly why an output is unsafe
Work you might do
- Reviewing model reasoning for clinically dangerous confidence
- Designing cases where the plausible answer is the wrong one
- Defining what a safe recommendation requires
- Flagging outputs that satisfy a rubric and fail a patient
T3 · Advanced STEM
01MathematicianPhD or equivalent
You will design problems where the final answer is verifiable and the path to it is not — and define what separates a correct derivation from a correct guess.
What we are looking for
- Academic or Research Depth — PhD in mathematics, or equivalent research output
- Proof Discipline — You can tell a proof from a plausible argument
- Problem Design — Ability to construct problems with unambiguous ground truth
- Clarity — Comfort writing for an intelligent non-specialist
Work you might do
- Designing problems where the obvious method fails
- Reviewing model derivations for gaps that do not change the answer
- Defining step-level correctness criteria
- Building problem sets with controlled difficulty
02PhysicistPhD or equivalent
You will define where a model's physical reasoning is dimensionally correct and physically impossible, and design tasks that separate the two.
What we are looking for
- Research Depth — PhD in physics, or equivalent research output
- Physical Intuition — You notice when a result violates something no equation stated
- Modelling Judgement — Comfort with approximation and its limits
- Precision — Ability to specify what makes a result valid
Work you might do
- Designing problems with realistic constraints and unambiguous answers
- Reviewing model reasoning for unphysical intermediate steps
- Defining validity checks beyond the final number
- Building graded problem sets across subfields
03ChemistPhD or equivalent
You will define where a proposed route, mechanism, or condition is theoretically sound and practically impossible.
What we are looking for
- Research Depth — PhD in chemistry, or equivalent industrial experience
- Bench Judgement — You know which reactions work on paper and not in a flask
- Safety Awareness — Comfort flagging hazardous or infeasible proposals
- Precision — Ability to specify exactly why a route fails
Work you might do
- Reviewing model-proposed syntheses for practical feasibility
- Designing tasks where the textbook answer is wrong
- Defining validity criteria for mechanisms and conditions
- Building problem sets with verifiable outcomes
04Biologist / Life SciencesPhD or equivalent
You will define where a biological claim is supported, where it is extrapolated, and where a model has produced a plausible sentence with no evidential basis.
What we are looking for
- Research Depth — PhD in a life science, or equivalent research output
- Evidence Discipline — You distinguish a finding from a hypothesis
- Method Literacy — Comfort assessing experimental design and statistics
- Clarity — Ability to state what a study does not show
Work you might do
- Reviewing model reasoning for unsupported biological claims
- Designing analysis tasks with verifiable conclusions
- Defining criteria for adequate evidence
- Flagging conclusions that outrun the data
05University ProfessorTop-ranked institution
You will help transform raw model behaviour into structured, teachable insight — writing and refining reasoning examples that reflect how experts actually argue.
What we are looking for
- Academic Standing — Current or former professor at a top-ranked institution, with published research
- Scholarly Depth — Advanced domain knowledge and the ability to explain it plainly
- Analytical Writing — Tightly reasoned, well-structured prose
- Precision — A methodical approach to critique and review
Work you might do
- Reviewing model responses for sound argumentation and conceptual clarity
- Drafting exemplar answers that illustrate expert reasoning
- Conducting comparative evaluations to surface subtle reasoning flaws
- Designing datasets that capture authentic expert cognition
T4 · Engineering & Machine Learning
01Post-Training ResearcherFrontier lab or equivalent
You will design reward functions and then try to break them. This is the work the labs say they are short of: writing a grader, attacking it with a model that is actively trying to cheat, and confirming afterwards that it measured what you intended.
What we are looking for
- Post-Training Experience — RLHF, RLVR, or reward-model work at a lab or serious research group
- Adversarial Mindset — You reach for the shortcut before you reach for the solution
- Systems Judgement — Understanding of how a weak reward propagates through a training run
- Rigour — Comfort with variance, pass^k, and contamination checks
Work you might do
- Designing invariant-based reward functions for multi-step tasks
- Running degenerate policies against graders and documenting the bypasses
- Reviewing reward specifications for exploitable slack
- Formalising what a robust verifier looks like in a new domain
02Machine Learning EngineerKaggle Grandmaster or Master preferred
You will identify the gap between a validation score and a real one. Leakage, split design, and the distance between what a model reports about itself and what the holdout says.
What we are looking for
- Competition Record — Kaggle Grandmaster or Master, or equivalent applied track record
- Leakage Instinct — You find the target buried in a feature before you trust the score
- Split Discipline — Time-aware, group-aware, and distribution-shift-aware validation
- Honesty — Comfort reporting that your own result did not hold
Work you might do
- Designing tasks where the obvious validation strategy is the wrong one
- Reviewing agent runs for leakage, tampered folds, and unfixed seeds
- Measuring the gap between self-reported CV and official holdout
- Building submissions that score well for the wrong reasons
03Software EngineerProduction experience, any stack
You will review code an agent wrote and decide whether it works, whether it would survive review, and whether the tests that passed actually tested anything.
What we are looking for
- Shipping Experience — You have maintained code other people depended on
- Review Judgement — You can tell a passing test from a meaningful one
- Debugging Skill — Finding the bug rather than the symptom
- Clarity — Writing a review comment that explains the failure
Work you might do
- Reviewing and ranking model-generated solutions with technical rationale
- Writing reference implementations a model should have produced
- Designing tests that catch the edge cases a model missed
- Flagging patches that pass CI and break production
04Data ScientistApplied, production experience
You will build the tasks where feature engineering quietly decides the outcome, and define what separates a legitimate feature from one that will not exist at inference time.
What we are looking for
- Production Experience — You have shipped a model other people depended on
- Feature Discipline — Clear judgement on what is available at prediction time
- Diagnostic Skill — Distribution shift, drift, and silent degradation
- Clarity — Ability to explain why a metric improved for the wrong reason
Work you might do
- Designing datasets with deliberately planted leakage
- Reviewing agent feature engineering for future information
- Measuring distribution-shift sensitivity across tasks
- Writing the rationale behind each planted trap
T5 · Operations
01Treasury OperationsCorporate or bank treasury
You will define what makes cash actually available — settlement timing, nostro balances, and the difference between a balance on a screen and money you can move.
What we are looking for
- Treasury Experience — Cash management, liquidity, or funding operations
- Settlement Literacy — T+N mechanics and the cost of assuming otherwise
- Operational Judgement — You have handled a funding gap in real time
- Precision — Comfort specifying availability rules exactly
Work you might do
- Defining availability rules that respect settlement cycles
- Reviewing agent runs where unsettled cash was treated as available
- Designing liquidity-constrained scenarios with real funding pressure
- Formalising nostro and suspense reconciliation
02Supply Chain PlannerManufacturing, distribution, or 3PL
You will define the constraints that make a plan executable rather than merely feasible on paper — lead times, capacity, and the ways a schedule satisfies the model and fails in the warehouse.
What we are looking for
- Planning Experience — S&OP, production planning, or distribution planning ownership
- Constraint Realism — You know which constraints are soft in a spreadsheet and hard in a building
- Systems Familiarity — ERP or planning system experience
- Clarity — Ability to state why a plan will not survive contact
Work you might do
- Defining executable-plan constraints for agent scheduling tasks
- Reviewing agent plans for violations a naive checker would pass
- Designing disruption scenarios with realistic recovery options
- Formalising what a feasible plan means
03Customs & Trade ComplianceLicensed broker or in-house
You will encode classification and duty logic — where the rules are objective, where they are contested, and where a defensible answer differs from a correct one.
What we are looking for
- Classification Experience — Licensed customs broker or in-house trade compliance
- Regulatory Currency — You track how the rules changed this year
- Judgement Under Ambiguity — Comfort with contested classifications
- Documentation Discipline — Every determination has a cited basis
Work you might do
- Defining classification tasks with defensible ground truth
- Reviewing agent determinations for unsupported reasoning
- Designing scenarios where the obvious classification is wrong
- Formalising duty and origin calculation checks
04Cross-border E-commerce Operations1688, Shopee, Lazada, or equivalent
You will define the workflows on the platforms nobody has simulated — the ones most of the world's operators actually open every morning.
What we are looking for
- Platform Experience — Hands-on seller or ops experience on non-US marketplaces
- Process Knowledge — Listing, pricing, fulfilment, dispute, and returns flows
- Local Fluency — You know the rules that are never written in English
- Precision — Comfort turning a habit into a specification
Work you might do
- Documenting real platform workflows step by step
- Reviewing agent runs for actions that are technically valid and operationally wrong
- Designing tasks with realistic platform constraints
- Formalising what a completed listing or dispute looks like
05Travel & Itinerary OperationsCorporate travel, DMC, or tour ops
You will define the scheduling constraints that hold in practice and not on paper — last admission times, weekly closures, queue times, and the buffers a plan needs to survive one delay.
What we are looking for
- Operations Experience — Corporate travel, destination management, or tour operations
- Local Knowledge — Asian markets preferred: closing days, reservation windows, transit realities
- Constraint Realism — You know which itineraries look fine and fall apart on the day
- Precision — Comfort turning local practice into an explicit rule
Work you might do
- Defining schedule constraints that chain: last admission, queue, and transit together
- Reviewing agent itineraries for plans feasible only on paper
- Designing tasks with realistic closure and reservation rules
- Formalising what a workable itinerary means
T6 · General
01Generalist ReviewerDegree-level, fluent English
You will review agent runs and flag where the output looks right and the work was not done. No domain licence required — careful reading and clear writing are the qualification.
What we are looking for
- Careful Reading — You notice when a conclusion does not follow from what came before
- Clear Writing — Complete sentences, specific objections
- Consistency — Comfort applying the same standard across many runs
- Honesty — Willingness to say you are unsure
Work you might do
- Reviewing agent trajectories against a written standard
- Flagging outputs that satisfy the format and miss the task
- Writing short, specific notes on why a run failed
- Helping calibrate difficulty across a task set
02Multilingual SpecialistChinese, Japanese, Korean, Thai, Indonesian
You will work on the software stack most of the world uses and nobody has simulated. Feishu, DingTalk, WeCom, Kingdee, Yonyou, 1688, Shopee, Lazada, LINE, KakaoTalk.
What we are looking for
- Native Fluency — In at least one of the listed languages, plus working English
- Platform Familiarity — Real use of the local enterprise or commerce tools
- Cultural Precision — You catch what a translation loses
- Care — Comfort with detailed, repetitive review
Work you might do
- Documenting workflows on non-US enterprise software
- Reviewing agent runs for locale-specific failures
- Translating task specifications without losing constraints
- Flagging conventions a non-local reviewer would miss
Apply
One question decides most of this.
The form takes about a minute. The last question is the one we actually read.
Being approximately right is the same as being wrong.