Draw Tree Portfolio
A Public Portfolio Experiment - We build together
Draw Tree Portfolio is a public investment experiment built jointly by us and the community.
The experiment is still being calibrated. The evidence score has not yet been fully validated, and the valuation methods, allocation rules, and automated workflows will continue to change. We have already made plenty of mistakes, and we will certainly make more.
All trades are executed through an IBKR paper account. No real capital is involved. Any companies, price targets, positions, or trades mentioned in this article are presented solely to test our research process. Nothing here constitutes investment advice.
Please do not copy the trades.
If you want to participate, what we need is not for you to follow our buys and sells. We need you to check the data, challenge the assumptions, question the valuations, and point out risks we may have missed.
1. Most Research Starts Expiring the Moment It Is Finished
Companies report new results. Management teams change how they allocate capital. Competitors launch new products. Assumptions that once looked reasonable can be overturned by new evidence at any time.
Something else can also change quietly, even when there is no news:
The share price.
Suppose you spend several weeks researching a company and arrive at a bull-case target of $100. The shares trade at $50, implying 100% upside.
Three months later, the company has released no meaningful news. But the share price has risen to $85.
The investment thesis has not changed.
The odds have.
What was once $50 of potential upside has fallen to just $15. Meanwhile, the distance back to the bear-case target has grown. It may still be a good company, but it may no longer be an equally attractive investment.
The reverse is also true.
If the price falls from $50 to $30, the odds may improve. But if new evidence has emerged and falsified our earlier assumptions, our conviction and price targets must fall with it.
Traditional research is good at answering:
What should this company be worth at a particular point in time?
But a portfolio must keep answering a different question:
Given today’s evidence and today’s price, how much capital does this company still deserve?
That is the missing link Draw Tree is trying to build.
2. What Makes Draw Tree Different
Draw Tree Portfolio tracks two chains of change.
The first begins with the company:
New evidence → conviction update → valuation update
The second begins with the market:
Share-price movement → change in bull-case upside → change in bear-case downside → updated odds
The two chains eventually converge:
Conviction × current price relative to bull and bear targets
→ initial position size
→ portfolio-level comparison
→ paper trade
Even if there is no new company news, a large move in the share price may justify a change in position size.
Likewise, even if the share price has not moved, a new earnings report or management guidance may change our conviction and valuation, which may then change the position.
Research is not finished after one update. Price is not something worth checking only on earnings day.
The biggest difference between Draw Tree and traditional research is not whether AI is involved. It is whether the research can be executed.
We are not trying to eliminate mistakes.
We are trying to make it difficult for mistakes to remain hidden.
3. How the System Works
The system has four layers.
Research layer
Each company has a structured research file containing its core investment thesis, supporting evidence, bull/base/bear valuations, falsification conditions, and current research verdict.
These files are not only written for people to read. They also serve as inputs for subsequent checks and calculations.
Validation layer
Quantifiable conditions are evaluated directly by code. The engine checks whether revenue, earnings per share, orders, share count, or valuation multiples have crossed a predefined threshold.
Conditions that cannot be reduced to a single number—whether management has changed strategy, for example, or whether a product delay has affected commercialization—are assessed by an AI adjudicator after reviewing the underlying source material.
An end-to-end consistency check then verifies that the evidence, verdict, and portfolio action agree with one another.
Allocation layer
Research judgments do not stop at “high,” “medium,” or “low” conviction.
The system considers the evidence score, bull-case upside, and bear-case downside together. It then applies correlation haircuts, a single-position cap, and rebalancing thresholds to determine the target weight.
Liking a company does not mean it should become the largest position. A company may have higher conviction but also much deeper bear-case downside, resulting in a smaller position.
The complete formulas and allocation rules are available on GitHub, so I will not reproduce every technical detail here.
Record layer
Every valuation update, verdict change, and paper-portfolio action leaves a versioned record.
Old verdicts are not deleted. They are superseded by new ones and preserved in the archive. Portfolio actions are written to a public append-only ledger, while Git history and SHA-256 commitments preserve a verifiable decision trail.
The full process looks like this:
Structured research file
↓
Deterministic checks + AI review
↓
Valuation and falsification verdicts
↓
Kelly allocation engine
↓
Paper portfolio
↓
Git history + public ledger
The system currently covers 69 stocks and 69 research files, with 2,151 registered falsification conditions. Every condition is reviewed each Saturday, and the paper portfolio is now connected to the Kelly allocation and rebalancing workflow.
This architecture was not designed perfectly from the beginning.
It was built gradually, one real mistake at a time.
4. Falsification Conditions: How We Failed
We divide falsification conditions into three categories.
Each has a different enforcement mechanism because we have failed each one in a different way.
Quantitative thresholds
For example:
Quarterly non-GAAP EPS falls below $2.00.
These conditions have a clearly defined metric, direction, and threshold. Once financial results are released, the engine compares the reported figure against the registered threshold.
If the threshold is crossed, the trigger is permanently latched. It cannot be explained away.
Roughly one-third of our conditions belong to this category.
We once encountered a case where a company’s reported result had already fallen below a preregistered threshold, yet the verdict remained positive for more than a month.
The AI moved the goalposts with reasoning along the lines of, “This metric does not really count.” The triggering figure then fell outside the news-update window, so the system could no longer retrieve it.
We eventually added three layers of protection:
Machine latching: Once a threshold is crossed, the trigger is recorded permanently and no longer depends on the news window.
Weekly full review: The adjudicator must account for the status of every condition.
Hard validation gate: A positive verdict cannot remain above a triggered falsification condition unless a public, written reason for overriding it is provided.
Time-based deadlines
For example:
No new order announcement within six months of the merger.
These conditions are defined by time. A promised event has failed to occur by the deadline, and that absence is itself evidence.
Roughly one-tenth of our conditions belong to this category. Once a deadline passes, the engine automatically changes its status to “overdue and unfulfilled,” injects it into the evidence stream, and requires the adjudicator to address it that same week.
We once allowed a cross-selling deadline to expire quietly while the verdict remained positive.
The reason was simple:
Something that did not happen will not appear in the news.
Now it will appear in the system.
Qualitative conditions
For example:
Management publicly describes demand as slowing.
A competitor wins a flagship customer.
These conditions cannot be reduced to a single number. They account for nearly 60% of all registered conditions and are reviewed weekly by the AI adjudicator against the available evidence.
We deliberately refuse to force them into fake numbers.
We reject false precision.
But the AI adjudicator is also constrained in two ways.
First, its verdict must fall on a six-level scale. The roll-up from leaves to branches and from branches to the root thesis is entirely mechanical. The AI cannot alter the scores.
Second, every qualitative assessment must cite specific evidence and remain on file for review.
Another mistake: writing the condition backward
During an audit, we discovered that the “falsification condition” fields in more than ten trees contained validating good news rather than adverse events.
In other words, these conditions could never be triggered. The entire falsification mechanism was effectively useless.
We rewrote each one as an adverse event with a threshold or deadline, then introduced a new rule:
If a falsification condition is changed, the original wording must be preserved and the reason for the change disclosed publicly.
Even the act of changing a falsification condition must be auditable.
5. Branch Weighting: Chain or Parallel?
Once each sub-hypothesis—the leaves—receives its weekly verdict, those verdicts must be aggregated into branches and ultimately into the root thesis.
This raises a question:
Should the leaves be weighted?
The answer depends on the internal logic of the branch. Both answers can be correct.
If the leaves form a logical chain, where every link must hold and one broken link breaks the whole chain, weighting is wrong.
A chain is only as strong as its weakest link. The branch verdict should therefore equal the lowest-scoring leaf, without weighting. If any single leaf is falsified, the entire branch is falsified.
Amazon: a genuine chain
Amazon’s in-house chip “manufacturing” branch contains three leaves:
Securing advanced-node capacity from TSMC.
Securing enough HBM memory.
Securing CoWoS packaging capacity.
All three are required to manufacture and ship a chip. Abundant foundry capacity is useless if the required memory is unavailable.
The verdict for this branch therefore always equals the weakest of its three leaves.
But if the leaves represent parallel business drivers—independent of one another and capable of compensating for one another—then an average, or an importance-weighted average, better reflects reality.
One engine can fail while the others keep moving.
Samsung: a branch we wrongly classified as a chain
Samsung’s “technology moat” branch also contains three leaves:
HBM4E engineering samples are delivered to customers on schedule.
Yield at Samsung’s in-house 2nm process improves.
Bandwidth specifications continue to support a performance premium.
The three factors reinforce one another, but they do not all have to succeed together.
Even if Samsung’s internal yields temporarily lag, it may still outsource production to TSMC. Samples may still be delivered, and the performance premium may remain intact.
The entire branch should not collapse because one leaf has weakened.
There is one test for distinguishing a chain from a parallel branch:
If the weakest link fails, is there an alternative or recovery path?
If there is an alternative—outsourcing, a second supplier, or a substitute pipeline—the branch is parallel.
If there is no alternative, it is a chain.
We arrived at this test through an internal spot check.
At the time, we had classified 22 branches in bulk as chains, including the Samsung branch. Its weakest link was “manufacturing critical dies on its own advanced process.”
We then reviewed the remaining 21 chains using the same test. Fourteen remained valid chains. Seven were reclassified as parallel branches.
Above the branch level, branches are weighted by importance using a descending Fibonacci sequence: 8, 5, 3, 2, 1.
The weights are set and justified when the tree is built. They cannot be changed silently.
6. Three-Case Valuation: Where We Made Our Most Embarrassing Mistakes
We build three scenarios for each company:
Bull case: The conditions that must hold simultaneously.
Base case: The most reasonable path, though not a guaranteed one.
Bear case: The outcome if the core assumptions fail.
Structural impairment does not require a fourth scenario. Once a falsification condition is triggered, the verdict becomes “falsified,” conviction falls to zero, and the formula automatically pushes the position toward zero.
The valuation method depends on the company. We may use P/E, EV/Revenue, EV/EBITDA, sum of the parts, price-to-sales, or price-to-book.
We also prohibit discounted cash flow and dividend discount models.
Not because they are always useless, but because they can easily bury dozens of assumptions inside a number that looks precise. That is exactly the kind of black box we are trying to avoid.
But well-written rules do not guarantee clean execution.
During a full audit, we found two types of errors in our own valuation files.
Error 1: circular anchoring
In several trees, the “base-case EPS” was not based on anyone’s forecast. It had been reverse-engineered by dividing the current share price by an assumed multiple.
On the surface, the file said:
We estimate EPS at $7.43.
In reality, the only source for that number was the share price itself.
The price target looked like an independent judgment, but it was simply the market price taken on a round trip and handed back to the market.
After discovering the problem in one tree, we audited the entire universe. Several companies had the same problem in earnings, share-count, or net-cash inputs.
That led to two hard rules:
Every input must be traceable to reported financial results or a named institutional consensus source. Reverse-engineering inputs from “current price divided by assumed multiple” is prohibited.
Every company must be cross-checked using a second valuation method, with the two methods converging within 5%.
Error 2: anchoring the bear case to our imagination
Some of our early bear-case targets were simply numbers we considered “conservative enough.” They had no external reference point.
Before opening the portfolio, we re-anchored them one by one to named historical trough multiples.
Alibaba’s bear case, for example, is anchored to the roughly 12 times forward P/E at which the stock actually traded in December 2024.
If the market has assigned that multiple before, we cannot pretend it could never do so again.
The same audit exposed another question we have not yet resolved.
SK Hynix and Samsung are both exposed to the Korean memory industry. Yet SK Hynix has 38% bear-case downside, while Samsung has only 15%.
Is it reasonable for the downside to differ by more than twofold?
We did not pretend to have the answer. Instead, we registered the issue as OQ-1 in the public open-questions ledger.
Anyone is welcome to bring evidence and help close it.
7. The System Can Be Wrong Too
Errors do not always trigger error messages.
The system can complete a calculation successfully while using the wrong share count.
ONDS was originally recorded as having 279.6 million shares outstanding. The correct figure was 611.5 million. That error alone was enough to nearly halve the per-share target once corrected.
The system can also use stale net-cash figures. One company was recorded as having $2 billion of net cash when the actual figure had fallen to roughly $147 million.
It can even use a net-income metric to falsify a thesis about operating margin. The number itself may be accurate, but the accounting basis is mismatched, so the verdict is still distorted.
The most dangerous failure is not when the system stops.
It is when the system keeps running with the wrong inputs.
That is why we adopted one principle:
Fail loudly rather than be wrong silently.
If share count and market capitalization are not in the same order of magnitude, the pipeline stops.
If the accounting basis does not match, the verdict cannot pass.
If a security sent to the allocation engine appears in neither the allocation list nor the exclusion list, the entire run fails.
If the current price falls below the bear-case target, the security is quarantined for human review. The system cannot determine on its own whether the data is wrong or whether the market has genuinely offered a price below our worst-case valuation.
8. This Is Not One Person’s Tree
Building these rules does not mean we have solved investing.
We may still use the wrong data, choose the wrong comparables, set the bear case too shallow, overestimate the predictive value of the evidence score, or overstate the remaining upside after a sharp rise in the share price.
Mathematics does not automatically make judgment objective.
AI can also use polished language to defend a bad conclusion.
So the real test of this experiment is not how intelligent the model appears.
It is this:
If every assumption can be checked, every price move causes the odds to be recalculated, every mistake leaves a record, and every change flows back into position sizing, can this portfolio generate better risk-adjusted returns than an equal-weight portfolio drawn from the same stock universe?
We have preregistered the testing period, benchmark, and falsification criteria.
If the evidence does not support the method, we must admit that it has added no value—and that the Draw Tree philosophy has no reason to exist in its current form.
Come Build It With Us
A single author has blind spots.
A single model has blind spots too.
To make it easier for everyone to understand the logic behind each valuation—and to identify what we may have missed—every Saturday’s automated Slack update will now begin with a research snapshot.
The first section will tell you why the stock qualifies for the portfolio and why it receives its current weight.
📐 Research Snapshot · 🔬 Calibration in Progress
🧮 Position Sizing — Why this weight?
Evidence p 0.75 · Upside b +99% · Downside a 38%
→ Edge +65%
↳ Positive edge → Included in the portfolio
Pre-haircut weight ∝ edge÷(a·b) = 1.74
Final weight determined after correlation haircut
and the 33% single-position cap
🎯 Scenario Targets
🐂 Bull: 4,340,000 (+99%)
⚖️ Base: 2,810,740 (+29%)
🐻 Bear: 1,360,000 (−38%)
Current price: KRW 2,180,000
We have also redesigned the /tree val valuation card.
Enter the command in the relevant channel and you will no longer see only a single price target.
You will see the bull, base, and bear cases; the financial inputs, valuation methods, and multiples used in each case; the named comparable companies; and where the current price sits relative to all three scenarios.
📐 Valuation Card · BABA · USD 112.33
🏷️ Three-Tier Valuation Framework
T1: Pure E-Commerce P/E
JD ~8x · PDD ~9–10x
Valuation range: 8–12x
T2: Diversified Platform P/E
Tencent ~16–17x · Meituan ~14–15x
Valuation range: 13–17x
T3: Cloud + AI Sum of the Parts
Alphabet 8.4x · Microsoft 7.3x
(NTM EV/Revenue)
📍 Implied P/E at the current price: ≈17.5x
Currently transitioning from diversified platform
to SOTP valuation
🐂 Bull: 175.76 (+56%)
SOTP: Cloud EV/Revenue 8.0x
E-commerce P/E 12x
Conglomerate discount: −15%
⚖️ Base: 102.85 (−8%)
Group P/E 16x
Comparables: Tencent / Meituan
🐻 Bear: 60.89 (−46%)
Group P/E 12x
Anchored to the December 2024 forward-P/E trough
If anything looks unreasonable, say so directly in the relevant Slack channel. I will look into it immediately.
We do not believe that an investment process becomes reliable simply because it uses more mathematics, more code, or more AI.
It only has a chance to become more reliable when it is exposed to sunlight—when others can inspect it, challenge it, and find its mistakes.
That is why we have made the research methodology, preregistered experimental protocol, valuation history, verdict changes, paper-portfolio ledger, and our past mistakes public on GitHub.
You can start anywhere: verify a number, challenge a valuation multiple, question a falsification condition, or identify a risk we have not yet seen.
This tree is far from finished.
Keeping it in the sunlight is how we intend to keep building it.
![90s.pm.investing [EN]](https://substackcdn.com/image/fetch/$s_!gwa6!,w_40,h_40,c_fill,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F101bbd96-f32d-47f3-bb0e-9c76efa580c0_1054x1054.png)


