Astrid Foundation Research workspace

Part 4 — The laboratory

14. Cell models — why skin becomes neurons

The core problem: STAG2 matters most in the brain, during development. You cannot biopsy a living child's brain.

The workaround rests on one fact: every cell in the body carries the same DNA, including the duplication. A skin cell and a neuron differ not in their DNA but in which genes are switched on. So you take cells you can safely obtain, and convert them into cells you can't.

The model hierarchy — each rung is more realistic and more expensive:

Model What it is Realism Cost/effort
Blood / LCL Immortalized white blood cells Low — wrong tissue entirely Very low
Fibroblasts Skin cells from a biopsy Low–moderate — wrong tissue, but real patient cells Low
iPSCs Reprogrammed stem cells Not a target tissue, but the gateway to everything Moderate
iPSC-derived neurons Actual patient neurons in a dish High — the right cell type High
Brain organoids 3D self-organizing mini-brain tissue Highest — some real architecture Very high
Mouse model Whole living animal Whole-body, but wrong species Very high

The standard path. A small skin biopsy yields fibroblasts, which are reprogrammed to iPSCs, which are then differentiated into neurons. Start to finish, biopsy to usable neurons typically takes something like six months to a year, and longer is common rather than alarming. If a programme you are following is on roughly that trajectory, it is on a normal timeline.

Two properties of iPSCs that change the strategic picture:

  1. They are effectively immortal. Once a good line exists it can be frozen, thawed, shipped, and expanded indefinitely. A child will not need a second biopsy for this research. One skin sample, unlimited experiments, forever.
  2. They are shareable. That same line can go to other labs worldwide. More labs working on the same cells = more shots on goal. This is a deliberate strategic lever, and it's worth negotiating consciously (Module 30).

A caution about the earlier evidence. The Kumar 2015 findings (Module 7) came from patient-derived cells that were not neurons. Findings in fibroblasts or blood cells are suggestive but not authoritative for a brain disease. This is the entire reason the neuron work matters — and the reason the first question about any screen is "does the signature replicate in the right cell type?"

Key terms:

References:


15. Reprogramming and differentiation — how it actually works

flowchart LR
    A["<b>Skin biopsy</b><br/><i>3mm punch</i>"] -->|"2–4 weeks<br/>outgrowth"| B["<b>Fibroblasts</b><br/>expanded in culture"]
    B -->|"<b>Reprogramming</b><br/>Yamanaka factors<br/>2–4 months"| C["<b>iPSCs</b><br/>quality-controlled line"]
    C -->|"<b>Differentiation</b><br/>1–3 months"| D["<b>Neurons</b><br/>in culture"]
    D --> E["<b>Phenotyping</b><br/>find a measurable difference<br/>3–12 months, highly variable"]
    E --> F["<b>Drug screen</b><br/>6–18 months"]

Reprogramming (fibroblast → iPSC). Four genes — the Yamanaka factors (OCT4, SOX2, KLF4, c-MYC) — are introduced into the skin cell. They erase its identity as a skin cell and return it to an embryo-like, pluripotent state, while leaving the genome intact. Shinya Yamanaka won the 2012 Nobel Prize for this.

It is finicky and frequently fails. Lines are lost, batches don't take, and months disappear. When a lab reports that reprogramming has succeeded, that is genuine, hard-won progress — not a routine box ticked.

Quality control that must happen — and is worth asking about. A newly made iPSC line is not automatically usable:

Check Why Question to ask
Karyotype / CNV array Reprogramming can introduce chromosomal damage. A damaged line produces garbage results. "Has the line been karyotyped?"
Pluripotency markers Confirms it really is pluripotent. "Confirmed pluripotent?"
Duplication still present & intact Confirms the disease-causing variant survived reprogramming. "Has the Xq25 duplication been re-confirmed in the iPSC line?"
Sterility / mycoplasma Contamination silently ruins experiments. "Mycoplasma-free?"
Identity / STR matching Confirms the line came from the intended donor — sample mix-ups are a real and documented problem. "Identity-matched to the original sample?"

These are all standard. A good lab does them as a matter of course, and will answer immediately and without defensiveness. Asking is not an accusation — it's fluency, and it's how you signal you'll be a serious partner.

Differentiation (iPSC → neuron). Two broad approaches, and the choice has real consequences:

Approach How Pros Cons
NGN2 induction Force-express the NEUROGENIN-2 gene; cells become neurons fast Fast (~2–3 weeks), very consistent, excellent for screening Somewhat artificial; produces a narrow neuron type; maturity limited
Dual-SMAD / patterned Mimic natural developmental signals step by step More faithful to real development; can specify regional identity Slower (2–3 months), more variable between batches

Why this matters for a chromatin disorder specifically. This condition is about gene regulation during development. A protocol that skips developmental steps (NGN2) may bypass the very window where the disease acts — potentially producing neurons that look deceptively normal. Conversely, patterned protocols are slower and more variable, which is painful for screening.

There is no universally right answer, and reasonable labs choose differently. But "which protocol, and why that one?" is a sharp, legitimate question, and the answer tells you a great deal about how carefully the experiment has been designed.

Key terms:

References:


16. Controls — the isogenic question

This is the most important methodological issue in the entire program, and the one most likely to be under-resourced.

To claim "the patient neurons are different," you must compare them to something. What you compare to determines whether the result means anything.

flowchart TD
    Q["The patient neurons differ from<br/>the control. Why?"]
    Q --> R1["Because of the<br/><b>STAG2 duplication</b><br/>✅ what we want to learn"]
    Q --> R2["Because of thousands of other<br/>genetic differences between<br/>two unrelated people<br/>❌ confounder"]
    Q --> R3["Because the lines were made<br/>or grown differently<br/>❌ batch effect"]

The control options, worst to best:

Control Strength Problem
Unrelated healthy donor Weak Differs from the patient at millions of positions. Any difference could be anything.
Several unrelated donors Moderate Averaging dilutes individual quirks — better, still not clean.
Parent / sibling Better Shares much of the genetic background.
Isogenic corrected line Gold standard Expensive and technically demanding.

What "isogenic" means. You take the patient's own iPSC line and use gene editing to remove the duplicated segment, creating a line that is genetically identical in every respect except the duplication. Now any difference between the two lines is attributable to the duplication and nothing else. It is the cleanest possible experiment.

Why it's hard for a duplication specifically. Correcting a spelling mistake is comparatively routine. Excising a ~443 kb duplicated segment cleanly requires cutting at two positions and having the cell rejoin the ends correctly — and it depends on the duplication's orientation, which brings us back to the breakpoint question from Module 3. It is not unprecedented: published work has demonstrated CRISPR removal of an X-chromosome duplication in MECP2-duplication patient cells, and single-guide excision of a tandem duplication in a DMD mouse model. So the approach sits within demonstrated technical competence — which is a genuine reason for cautious confidence.

It has been done in the sister disease. Rizvi et al., Molecular Therapy Nucleic Acids, December 2024 (PMID 39507402) engineered a human cell model carrying an IRAK1–MECP2 duplication, then used a single guide RNA CRISPR-Cas9 strategy to excise the duplicated segment, "notably halving both MECP2 and IRAK1 expression."

Be realistic about the yield all the same. Excision competes with other repair outcomes — the segment can invert rather than be removed — so this is a months-long undertaking with a low per-clone success rate, not a quick fix. But it is a technique this specific group has published, not a theoretical option.

💡 An elegant alternative that avoids CRISPR entirely. If the child's mother is a carrier, her cells can be reprogrammed to iPSCs and then sorted by which X chromosome is active (Module 2). Because X-inactivation is random, you can derive clones expressing the duplicated X and clones expressing the normal X — from the same person, same genome, differing only in which X is active. That is a genuinely isogenic pair, obtained by selection rather than editing. It is a standard published technique in Rett/MECP2 research.

This route is entirely contingent on maternal carrier status — which is one reason confirming it is worth doing early, and a good example of how a clinical question and a laboratory strategy can turn out to be the same question.

The pragmatic reality. Isogenic lines take months and meaningful money. A reasonable program often starts screening against unrelated controls and generates the isogenic line in parallel, using it to validate hits rather than to find them. That's a defensible sequencing of work — but it should be a deliberate, stated plan, not an omission.

The question to ask: "What are the patient neurons being compared against — and is an isogenic corrected line planned? If not now, at what stage?"

Key terms:

References:


17. Readouts — how you measure a difference

A readout (or assay) is the specific measurement that tells you whether cells are behaving abnormally, and whether a drug fixed it. Choosing it is the highest-leverage decision in the screen.

The central tension: the most meaningful measurements are usually the least scalable, and screening requires scale.

Readout What it measures Screen-compatible? Meaningfulness
qPCR / protein blot for STAG2 STAG2 level directly ✅ Yes, cheap ⚠️ Direct but shallow — proves engagement, not benefit
HiBiT reporter for STAG2 STAG2 protein level, as light ✅✅ Ideal for screening ⚠️ Direct but shallow; needs a counter-screen
OPHN1 expression The Kumar downstream marker ✅ Yes ✅ Disease-linked, and already published
RNA-seq signature Thousands of genes at once ⚠️ Costly at scale ✅✅ Richest; the true "fingerprint"
High-content imaging Neurite length, branching, nuclear shape ✅ Yes, well-suited ✅ Structural, disease-plausible
MEA electrophysiology Actual electrical firing and network activity ⚠️ Moderate throughput ✅✅ Closest to real neuronal function
Hi-C / CUT&RUN Chromatin architecture directly ❌ No — validation only ✅✅ Most mechanistically direct
Organoids 3D tissue-level phenotypes ❌ No — validation only ✅✅ Most realistic

The professional answer is a tiered funnel — cheap and scalable first, expensive and meaningful last:

flowchart TD
    A["<b>Tier 0 — in silico</b><br/>signature-reversal query<br/><i>thousands of compounds, ~free</i>"] --> B
    B["<b>Tier 1 — primary screen</b><br/>Cheap scalable readout on many compounds<br/><i>e.g. reporter, OPHN1, imaging</i>"] --> C
    C["<b>Tier 2 — confirmation</b><br/>Dose–response; is it real and dose-dependent?"] --> D
    D["<b>Tier 3 — orthogonal</b><br/>A <i>different</i> method must agree<br/><i>e.g. RNA-seq, MEA</i>"] --> E
    E["<b>Tier 4 — triage</b><br/>Brain penetrance? Pediatric safety?<br/>Approved? Reasonable dose?"] --> F
    F["<b>Tier 5 — validation</b><br/>Isogenic line, organoids,<br/>independent replication"]

💡 The readout most worth building: a HiBiT reporter. HiBiT is a tiny tag — just 11 amino acids — that can be attached to a protein inside the cell's own gene, so the cell makes STAG2 with a miniature light-emitting label permanently attached. Add a reagent and the well glows in proportion to how much STAG2 protein is present.

Why this is attractive: it reads the actual protein (not a proxy), needs no antibodies, gives an answer the same day, and works in 384-well plates — so it scales to thousands of compounds. There is a 2025 precedent using exactly this approach in patient iPSC-neurons for a different disease.

The mandatory companion is a counter-screen line — a second, unrelated gene tagged the same way. Without it, every compound that generally poisons transcription, translation, or the cell itself will make the light go down and look like a brilliant hit. This is not optional.

The single most important principle here: orthogonal confirmation. A hit found by one method, confirmed only by that same method, is not confirmed. It must be reproduced by a different kind of measurement. Screens produce false positives at high rates — this is normal and expected, not a sign of a bad lab — and orthogonal confirmation is the standard defense.

The phenotype problem — the real risk to be alert to. Everything above assumes there is a measurable, reproducible difference between the affected neurons and controls. If there isn't, there is nothing to screen against, and the drug screen cannot proceed. A beautiful model with no measurable phenotype is a genuine, common dead end in this field.

This is why "what phenotype have you found, and how reproducible is it?" is the question that matters most in any conversation about the model. Everything downstream is contingent on it.

Key terms:

References:


18. Drug screening and hit triage

What a screen is. Take many compounds, apply each to the cells, measure the readout, and look for compounds that move the abnormal measurement back toward normal. A compound that does is a hit.

Why repurposing, specifically. Instead of inventing a new molecule (10–15 years, hundreds of millions of dollars, high failure rate), you test drugs already approved for something else. Their human safety profile is already established — which is precisely what makes a fast path to a single child conceivable at all.

The named libraries you should know:

Library Approx. size Notes
Broad Drug Repurposing Hub ~6,000 Well-annotated; widely used in academia
NCATS Pharmaceutical Collection ~2,500–3,000 Approved drugs; NCATS also offers collaboration programs
ReFRAME (Calibr/Scripps) ~12,000 One of the most comprehensive; access via collaboration
Prestwick Chemical Library ~1,500 Off-patent, mostly approved; affordable and common
Selleck / MedChemExpress FDA-approved sets ~2,000–3,000 Commercially purchasable

Bigger is not automatically better. A larger library means more hits and more false positives, and a stronger statistical burden. For a small academic lab, a well-annotated ~1,500–6,000-compound approved-drug library is usually the sensible choice.

The best library may already be in the building. Many academic hospitals and universities run their own screening cores, typically holding standard collections — Prestwick, Microsource Spectrum, LOPAC, an ApexBio FDA-approved set (~1,500 compounds) and often a dedicated ApexBio epigenetics sub-library (~280 compounds) — available to in-house investigators on cost-recovery terms. That is frequently more accessible than the higher-profile collections, which carry real access constraints: ReFRAME's formal call for proposals is US-restricted, and Broad Hub plates come with data-sharing obligations. For a chromatin disorder, an epigenetics sub-library is especially relevant, and it is worth asking whether one is available locally before pursuing anything larger.

⚠️ The honest base rate — read this before hoping

There is no published compound that lowers the amount of STAG2. Not one. The STAG2-targeting small molecules in the literature do something different: StagX1 exploits STAG2 loss in Ewing sarcoma, and KPT-6566 binds STAG1 and STAG2 directly and blocks their interaction with the cohesin ring subunit SCC1/RAD21 (PMID 39541712). Both were built for cancer, where toxicity is the goal.

But note what KPT-6566 demonstrates: the STAG2–cohesin interaction is druggable, and there is a published high-throughput assay for finding compounds that hit it. That is a different and more tractable target than "lower the protein" — see the discussion of interaction-blocking below.

But be careful about a claim you will encounter. It is often said that the sister disease already ran this experiment — that a roughly 28,000-compound screen for MeCP2 modulators found nothing that lowered the protein. That is a misreading, and it matters because it is the single most discouraging thing said about this approach.

That screen (PMC7657357) was a Rett syndrome screen — aimed at the opposite problem, reactivating a silenced copy of the gene. Its published methods score only increases: "a twofold or more change in relative luciferase unit was considered an active compound." It never asked what lowers MeCP2, and reports nothing about compounds that did.

So the honest position is not "this was tried and it failed." It is nobody has properly looked — in either disease. What the MECP2-duplication field did do is move to oligonucleotides and RNA editing, which is a statement about where they placed their bets, not about a screen that came back empty.

That is the closest available base rate, and it is discouraging. It does not mean the screen shouldn't run — these neurons are a different system, the model has independent scientific value, and the cost of looking is modest. But it does mean three things:

  1. Expect the screen to fail more likely than not. Plan emotionally and strategically for that.
  2. Agree a kill criterion in advance — what result would mean "stop and pivot."
  3. Start the durable lane now, in parallel. If the small-molecule route is a long shot, the oligonucleotide route stops being the backup plan and becomes the main plan. It is also the slowest, which is exactly why it can't wait for the screen to finish.

This is the single most important expectation-setting passage in this primer. Knowing the base rate in advance is what lets you hear a negative screen result as information rather than as catastrophe.

Hit triage — the filters that matter before anyone gets excited:

  1. Is it real? Dose-dependent, reproducible, and confirmed by an orthogonal method.
  2. Is it specific, or just toxic? Many "hits" simply make cells sick, which changes every measurement. This is the most common false positive in the field. A proper screen always runs a parallel cytotoxicity/viability assay.
  3. Does it reach the brain? A drug that can't cross the blood-brain barrier cannot work here, no matter how good the dish data. Filter early.
  4. Is it safe in children, chronically? A child would potentially take this for years. An oncology drug with serious toxicity is a very different proposition from a well-tolerated approved medicine.
  5. Is the effective dose achievable in a person? A compound that works at a concentration unreachable in a human brain is not a candidate.
  6. What's the mechanism? Not strictly required, but a plausible mechanism greatly strengthens the case and guides dosing.

A caution about the emotional dynamics. The word "hit" sounds like a discovery. In screening it means "a signal worth checking." Most hits do not survive triage — that is the normal, expected, correctly-functioning process, not a failure. Knowing this in advance is genuinely protective: it lets you hear "we found some hits" as encouraging news about the process rather than as news about a treatment.

Key terms:

References: