
September 28, 2026
TypeSafe calls Jev the first model in its System One model class, named after Daniel Kahneman’s distinction between fast, automatic judgment (System 1) and slower, deliberate reasoning (System 2). Also, TypeSafe claims that System One models output calibrated probabilities. Calibrated against what? I’m not the first to ask, but I want to answer a more specific question. Kahneman himself later admitted that he had placed too much faith in underpowered studies. Can Jev give plausible probabilities for a question where published research has produced exaggerated estimates?
For example, does Jev “believe” the widely reported claim that good-looking parents are 36% more likely to have a daughter than a son as their first child? Gelman and Weakliem trace how this claim grew in analysis and reporting. A regression analysis estimated a statistically non-significant difference of 4.7 percentage points. A selected comparison put the difference at 8 percentage points; a logistic regression estimated 26% lower odds of having a son; and media reports ultimately turned this into the claim that attractive parents were 36% more likely to have a daughter than a son. The study had about 3,000 respondents, but even that sample was too small to reliably detect the effects of less than 1 percentage point considered plausible from prior research.
Let’s ask Jev a question using the same beauty-rating scale. We do not provide the original study or identify a particular population.
An interviewer rates the beauty of a random parent on a scale from 1 to 5. The rating of that parent is X. Is the parent’s first child a girl?
We sample 1,000 ratings uniformly between 1 and 5 and make a separate call to jev-1.13.0 for each sampled rating. Following TypeSafe’s self-consistency cookbook, every call contains one question and a fresh, throwaway uid in its state. Because the uid changes on every call, the inputs are not literally identical. If the answers vary, we cannot tell whether this reflects Jev’s inherent predictive variability or sensitivity to the otherwise irrelevant uid. Each dot below is the answer from one separate API call:

Points are jittered horizontally so we can get a better sense of the distribution. The average probabilities for ratings 1 through 5 are 43.9%, 45.0%, 46.3%, 44.6%, and 42.4%. All are below the usual baseline of about 49%. The largest difference between group averages is 3.9 percentage points, between ratings 3 and 5. The non-monotonic pattern is also odd: the average rises from 43.9% at rating 1 to 46.3% at rating 3, then falls to 42.4% at rating 5. Although the prompt leaves the population unspecified, these values are difficult to reconcile with that baseline and the small effects considered plausible in the literature.
While this experiment does not measure calibration against observed outcomes, it does show that Jev predicts probabilities that are difficult to reconcile with the usual birth-sex baseline and plausible effects of parental attractiveness. What evidence supports treating Jev’s probabilities as calibrated on questions like this?