Part 4 — The laboratory
14. Cell models — why skin becomes neurons
The core problem: STAG2 matters most in the brain, during development. You cannot biopsy a living child's brain.
The workaround rests on one fact: every cell in the body carries the same DNA, including the duplication. A skin cell and a neuron differ not in their DNA but in which genes are switched on. So you take cells you can safely obtain, and convert them into cells you can't.
The model hierarchy — each rung is more realistic and more expensive:
| Model | What it is | Realism | Cost/effort |
|---|---|---|---|
| Blood / LCL | Immortalized white blood cells | Low — wrong tissue entirely | Very low |
| Fibroblasts | Skin cells from a biopsy | Low–moderate — wrong tissue, but real patient cells | Low |
| iPSCs | Reprogrammed stem cells | Not a target tissue, but the gateway to everything | Moderate |
| iPSC-derived neurons | Actual patient neurons in a dish | High — the right cell type | High |
| Brain organoids | 3D self-organizing mini-brain tissue | Highest — some real architecture | Very high |
| Mouse model | Whole living animal | Whole-body, but wrong species | Very high |
The standard path. A small skin biopsy yields fibroblasts, which are reprogrammed to iPSCs, which are then differentiated into neurons. Start to finish, biopsy to usable neurons typically takes something like six months to a year, and longer is common rather than alarming. If a programme you are following is on roughly that trajectory, it is on a normal timeline.
Two properties of iPSCs that change the strategic picture:
- They are effectively immortal. Once a good line exists it can be frozen, thawed, shipped, and expanded indefinitely. A child will not need a second biopsy for this research. One skin sample, unlimited experiments, forever.
- They are shareable. That same line can go to other labs worldwide. More labs working on the same cells = more shots on goal. This is a deliberate strategic lever, and it's worth negotiating consciously (Module 30).
A caution about the earlier evidence. The Kumar 2015 findings (Module 7) came from patient-derived cells that were not neurons. Findings in fibroblasts or blood cells are suggestive but not authoritative for a brain disease. This is the entire reason the neuron work matters — and the reason the first question about any screen is "does the signature replicate in the right cell type?"
Key terms:
- Fibroblast — the connective-tissue cell grown from a skin biopsy.
- LCL (lymphoblastoid cell line) — immortalized patient blood cells; convenient, low realism.
- iPSC (induced pluripotent stem cell) — an adult cell reprogrammed to an embryo-like state.
- Pluripotent — able to become nearly any cell type.
- Differentiation — guiding a stem cell to become a specific cell type.
- Organoid — 3D self-organizing tissue grown from stem cells.
- Cell line — a population of cells maintained in culture.
- Passage — one round of splitting and re-plating cells; high passage number can mean drift.
References:
- NIH — Stem Cell Basics — authoritative introduction.
15. Reprogramming and differentiation — how it actually works
flowchart LR
A["<b>Skin biopsy</b><br/><i>3mm punch</i>"] -->|"2–4 weeks<br/>outgrowth"| B["<b>Fibroblasts</b><br/>expanded in culture"]
B -->|"<b>Reprogramming</b><br/>Yamanaka factors<br/>2–4 months"| C["<b>iPSCs</b><br/>quality-controlled line"]
C -->|"<b>Differentiation</b><br/>1–3 months"| D["<b>Neurons</b><br/>in culture"]
D --> E["<b>Phenotyping</b><br/>find a measurable difference<br/>3–12 months, highly variable"]
E --> F["<b>Drug screen</b><br/>6–18 months"]
Reprogramming (fibroblast → iPSC). Four genes — the Yamanaka factors (OCT4, SOX2, KLF4, c-MYC) — are introduced into the skin cell. They erase its identity as a skin cell and return it to an embryo-like, pluripotent state, while leaving the genome intact. Shinya Yamanaka won the 2012 Nobel Prize for this.
It is finicky and frequently fails. Lines are lost, batches don't take, and months disappear. When a lab reports that reprogramming has succeeded, that is genuine, hard-won progress — not a routine box ticked.
Quality control that must happen — and is worth asking about. A newly made iPSC line is not automatically usable:
| Check | Why | Question to ask |
|---|---|---|
| Karyotype / CNV array | Reprogramming can introduce chromosomal damage. A damaged line produces garbage results. | "Has the line been karyotyped?" |
| Pluripotency markers | Confirms it really is pluripotent. | "Confirmed pluripotent?" |
| Duplication still present & intact | Confirms the disease-causing variant survived reprogramming. | "Has the Xq25 duplication been re-confirmed in the iPSC line?" |
| Sterility / mycoplasma | Contamination silently ruins experiments. | "Mycoplasma-free?" |
| Identity / STR matching | Confirms the line came from the intended donor — sample mix-ups are a real and documented problem. | "Identity-matched to the original sample?" |
These are all standard. A good lab does them as a matter of course, and will answer immediately and without defensiveness. Asking is not an accusation — it's fluency, and it's how you signal you'll be a serious partner.
Differentiation (iPSC → neuron). Two broad approaches, and the choice has real consequences:
| Approach | How | Pros | Cons |
|---|---|---|---|
| NGN2 induction | Force-express the NEUROGENIN-2 gene; cells become neurons fast | Fast (~2–3 weeks), very consistent, excellent for screening | Somewhat artificial; produces a narrow neuron type; maturity limited |
| Dual-SMAD / patterned | Mimic natural developmental signals step by step | More faithful to real development; can specify regional identity | Slower (2–3 months), more variable between batches |
Why this matters for a chromatin disorder specifically. This condition is about gene regulation during development. A protocol that skips developmental steps (NGN2) may bypass the very window where the disease acts — potentially producing neurons that look deceptively normal. Conversely, patterned protocols are slower and more variable, which is painful for screening.
There is no universally right answer, and reasonable labs choose differently. But "which protocol, and why that one?" is a sharp, legitimate question, and the answer tells you a great deal about how carefully the experiment has been designed.
Key terms:
- Yamanaka factors — OCT4, SOX2, KLF4, c-MYC; the reprogramming cocktail.
- Karyotype — an image/analysis of chromosomes, checking for damage.
- NGN2 (Neurogenin-2) — a gene that rapidly forces neuronal identity.
- Dual-SMAD inhibition — a patterning method mimicking natural neural development.
- Cortical neuron — the neuron type of the cerebral cortex; most relevant here.
- Maturity — how developmentally advanced the neurons are; dish neurons are usually immature.
- Mycoplasma — a common, invisible bacterial contaminant of cultures.
- STR matching — a DNA fingerprint confirming cell identity.
References:
- PubMed search: NGN2 induced neurons iPSC protocol
- Dissecting transcriptomic signatures of neuronal differentiation and maturation using iPSCs — Nature Communications
16. Controls — the isogenic question
This is the most important methodological issue in the entire program, and the one most likely to be under-resourced.
To claim "the patient neurons are different," you must compare them to something. What you compare to determines whether the result means anything.
flowchart TD
Q["The patient neurons differ from<br/>the control. Why?"]
Q --> R1["Because of the<br/><b>STAG2 duplication</b><br/>✅ what we want to learn"]
Q --> R2["Because of thousands of other<br/>genetic differences between<br/>two unrelated people<br/>❌ confounder"]
Q --> R3["Because the lines were made<br/>or grown differently<br/>❌ batch effect"]
The control options, worst to best:
| Control | Strength | Problem |
|---|---|---|
| Unrelated healthy donor | Weak | Differs from the patient at millions of positions. Any difference could be anything. |
| Several unrelated donors | Moderate | Averaging dilutes individual quirks — better, still not clean. |
| Parent / sibling | Better | Shares much of the genetic background. |
| Isogenic corrected line | ✅ Gold standard | Expensive and technically demanding. |
What "isogenic" means. You take the patient's own iPSC line and use gene editing to remove the duplicated segment, creating a line that is genetically identical in every respect except the duplication. Now any difference between the two lines is attributable to the duplication and nothing else. It is the cleanest possible experiment.
Why it's hard for a duplication specifically. Correcting a spelling mistake is comparatively routine. Excising a ~443 kb duplicated segment cleanly requires cutting at two positions and having the cell rejoin the ends correctly — and it depends on the duplication's orientation, which brings us back to the breakpoint question from Module 3. It is not unprecedented: published work has demonstrated CRISPR removal of an X-chromosome duplication in MECP2-duplication patient cells, and single-guide excision of a tandem duplication in a DMD mouse model. So the approach sits within demonstrated technical competence — which is a genuine reason for cautious confidence.
It has been done in the sister disease. Rizvi et al., Molecular Therapy Nucleic Acids, December 2024 (PMID 39507402) engineered a human cell model carrying an IRAK1–MECP2 duplication, then used a single guide RNA CRISPR-Cas9 strategy to excise the duplicated segment, "notably halving both MECP2 and IRAK1 expression."
Be realistic about the yield all the same. Excision competes with other repair outcomes — the segment can invert rather than be removed — so this is a months-long undertaking with a low per-clone success rate, not a quick fix. But it is a technique this specific group has published, not a theoretical option.
💡 An elegant alternative that avoids CRISPR entirely. If the child's mother is a carrier, her cells can be reprogrammed to iPSCs and then sorted by which X chromosome is active (Module 2). Because X-inactivation is random, you can derive clones expressing the duplicated X and clones expressing the normal X — from the same person, same genome, differing only in which X is active. That is a genuinely isogenic pair, obtained by selection rather than editing. It is a standard published technique in Rett/MECP2 research.
This route is entirely contingent on maternal carrier status — which is one reason confirming it is worth doing early, and a good example of how a clinical question and a laboratory strategy can turn out to be the same question.
The pragmatic reality. Isogenic lines take months and meaningful money. A reasonable program often starts screening against unrelated controls and generates the isogenic line in parallel, using it to validate hits rather than to find them. That's a defensible sequencing of work — but it should be a deliberate, stated plan, not an omission.
The question to ask: "What are the patient neurons being compared against — and is an isogenic corrected line planned? If not now, at what stage?"
Key terms:
- Isogenic — genetically identical except for the one variant under study.
- Confounder — an alternative explanation for an observed difference.
- Batch effect — differences caused by when/how samples were processed rather than by biology.
- Genetic background — all the other variation in a person's genome.
- CRISPR/Cas9 — the gene-editing system used to make targeted DNA changes.
- Technical vs biological replicate — same sample measured repeatedly vs genuinely independent samples. Only biological replicates support strong claims.
References:
17. Readouts — how you measure a difference
A readout (or assay) is the specific measurement that tells you whether cells are behaving abnormally, and whether a drug fixed it. Choosing it is the highest-leverage decision in the screen.
The central tension: the most meaningful measurements are usually the least scalable, and screening requires scale.
| Readout | What it measures | Screen-compatible? | Meaningfulness |
|---|---|---|---|
| qPCR / protein blot for STAG2 | STAG2 level directly | ✅ Yes, cheap | ⚠️ Direct but shallow — proves engagement, not benefit |
| HiBiT reporter for STAG2 | STAG2 protein level, as light | ✅✅ Ideal for screening | ⚠️ Direct but shallow; needs a counter-screen |
| OPHN1 expression | The Kumar downstream marker | ✅ Yes | ✅ Disease-linked, and already published |
| RNA-seq signature | Thousands of genes at once | ⚠️ Costly at scale | ✅✅ Richest; the true "fingerprint" |
| High-content imaging | Neurite length, branching, nuclear shape | ✅ Yes, well-suited | ✅ Structural, disease-plausible |
| MEA electrophysiology | Actual electrical firing and network activity | ⚠️ Moderate throughput | ✅✅ Closest to real neuronal function |
| Hi-C / CUT&RUN | Chromatin architecture directly | ❌ No — validation only | ✅✅ Most mechanistically direct |
| Organoids | 3D tissue-level phenotypes | ❌ No — validation only | ✅✅ Most realistic |
The professional answer is a tiered funnel — cheap and scalable first, expensive and meaningful last:
flowchart TD
A["<b>Tier 0 — in silico</b><br/>signature-reversal query<br/><i>thousands of compounds, ~free</i>"] --> B
B["<b>Tier 1 — primary screen</b><br/>Cheap scalable readout on many compounds<br/><i>e.g. reporter, OPHN1, imaging</i>"] --> C
C["<b>Tier 2 — confirmation</b><br/>Dose–response; is it real and dose-dependent?"] --> D
D["<b>Tier 3 — orthogonal</b><br/>A <i>different</i> method must agree<br/><i>e.g. RNA-seq, MEA</i>"] --> E
E["<b>Tier 4 — triage</b><br/>Brain penetrance? Pediatric safety?<br/>Approved? Reasonable dose?"] --> F
F["<b>Tier 5 — validation</b><br/>Isogenic line, organoids,<br/>independent replication"]
💡 The readout most worth building: a HiBiT reporter. HiBiT is a tiny tag — just 11 amino acids — that can be attached to a protein inside the cell's own gene, so the cell makes STAG2 with a miniature light-emitting label permanently attached. Add a reagent and the well glows in proportion to how much STAG2 protein is present.
Why this is attractive: it reads the actual protein (not a proxy), needs no antibodies, gives an answer the same day, and works in 384-well plates — so it scales to thousands of compounds. There is a 2025 precedent using exactly this approach in patient iPSC-neurons for a different disease.
The mandatory companion is a counter-screen line — a second, unrelated gene tagged the same way. Without it, every compound that generally poisons transcription, translation, or the cell itself will make the light go down and look like a brilliant hit. This is not optional.
The single most important principle here: orthogonal confirmation. A hit found by one method, confirmed only by that same method, is not confirmed. It must be reproduced by a different kind of measurement. Screens produce false positives at high rates — this is normal and expected, not a sign of a bad lab — and orthogonal confirmation is the standard defense.
The phenotype problem — the real risk to be alert to. Everything above assumes there is a measurable, reproducible difference between the affected neurons and controls. If there isn't, there is nothing to screen against, and the drug screen cannot proceed. A beautiful model with no measurable phenotype is a genuine, common dead end in this field.
This is why "what phenotype have you found, and how reproducible is it?" is the question that matters most in any conversation about the model. Everything downstream is contingent on it.
Key terms:
- Assay / readout — the specific measurement made.
- qPCR — quantifies how much of a specific mRNA is present.
- Western blot — measures protein amount.
- RNA-seq — sequences all mRNA; gives a genome-wide expression profile.
- High-content imaging — automated microscopy plus image analysis at scale.
- MEA (multi-electrode array) — a plate of electrodes recording neuronal firing.
- Hi-C / Micro-C — methods that map 3D genome folding.
- CUT&RUN / ChIP-seq — map where a protein (e.g. cohesin, CTCF) sits on DNA.
- Throughput — how many conditions can be tested per unit time/cost.
- Orthogonal validation — confirmation by an independent method.
- Dose–response — more drug produces more effect; a hallmark of a real pharmacological effect.
- Z-factor — a statistic measuring whether an assay separates signal from noise well enough to screen (>0.5 is generally considered good).
- HiBiT — an 11-amino-acid luminescent tag attached to a protein inside the cell's own gene; the well glows in proportion to how much of that protein is present.
- Reporter line — a cell line engineered so that something you care about becomes easy to see or measure (here, light).
- Knock-in — inserting a sequence at a precise location in the genome (as opposed to a knock-out, which removes function).
- 384-well plate — a tray of 384 tiny wells; the standard unit of screening scale.
- Counter-screen — a parallel test designed to catch trivial explanations, especially "the compound is just poisoning the cells."
References:
18. Drug screening and hit triage
What a screen is. Take many compounds, apply each to the cells, measure the readout, and look for compounds that move the abnormal measurement back toward normal. A compound that does is a hit.
Why repurposing, specifically. Instead of inventing a new molecule (10–15 years, hundreds of millions of dollars, high failure rate), you test drugs already approved for something else. Their human safety profile is already established — which is precisely what makes a fast path to a single child conceivable at all.
The named libraries you should know:
| Library | Approx. size | Notes |
|---|---|---|
| Broad Drug Repurposing Hub | ~6,000 | Well-annotated; widely used in academia |
| NCATS Pharmaceutical Collection | ~2,500–3,000 | Approved drugs; NCATS also offers collaboration programs |
| ReFRAME (Calibr/Scripps) | ~12,000 | One of the most comprehensive; access via collaboration |
| Prestwick Chemical Library | ~1,500 | Off-patent, mostly approved; affordable and common |
| Selleck / MedChemExpress FDA-approved sets | ~2,000–3,000 | Commercially purchasable |
Bigger is not automatically better. A larger library means more hits and more false positives, and a stronger statistical burden. For a small academic lab, a well-annotated ~1,500–6,000-compound approved-drug library is usually the sensible choice.
The best library may already be in the building. Many academic hospitals and universities run their own screening cores, typically holding standard collections — Prestwick, Microsource Spectrum, LOPAC, an ApexBio FDA-approved set (~1,500 compounds) and often a dedicated ApexBio epigenetics sub-library (~280 compounds) — available to in-house investigators on cost-recovery terms. That is frequently more accessible than the higher-profile collections, which carry real access constraints: ReFRAME's formal call for proposals is US-restricted, and Broad Hub plates come with data-sharing obligations. For a chromatin disorder, an epigenetics sub-library is especially relevant, and it is worth asking whether one is available locally before pursuing anything larger.
⚠️ The honest base rate — read this before hoping
There is no published compound that lowers the amount of STAG2. Not one. The STAG2-targeting small molecules in the literature do something different: StagX1 exploits STAG2 loss in Ewing sarcoma, and KPT-6566 binds STAG1 and STAG2 directly and blocks their interaction with the cohesin ring subunit SCC1/RAD21 (PMID 39541712). Both were built for cancer, where toxicity is the goal.
But note what KPT-6566 demonstrates: the STAG2–cohesin interaction is druggable, and there is a published high-throughput assay for finding compounds that hit it. That is a different and more tractable target than "lower the protein" — see the discussion of interaction-blocking below.
But be careful about a claim you will encounter. It is often said that the sister disease already ran this experiment — that a roughly 28,000-compound screen for MeCP2 modulators found nothing that lowered the protein. That is a misreading, and it matters because it is the single most discouraging thing said about this approach.
That screen (PMC7657357) was a Rett syndrome screen — aimed at the opposite problem, reactivating a silenced copy of the gene. Its published methods score only increases: "a twofold or more change in relative luciferase unit was considered an active compound." It never asked what lowers MeCP2, and reports nothing about compounds that did.
So the honest position is not "this was tried and it failed." It is nobody has properly looked — in either disease. What the MECP2-duplication field did do is move to oligonucleotides and RNA editing, which is a statement about where they placed their bets, not about a screen that came back empty.
That is the closest available base rate, and it is discouraging. It does not mean the screen shouldn't run — these neurons are a different system, the model has independent scientific value, and the cost of looking is modest. But it does mean three things:
- Expect the screen to fail more likely than not. Plan emotionally and strategically for that.
- Agree a kill criterion in advance — what result would mean "stop and pivot."
- Start the durable lane now, in parallel. If the small-molecule route is a long shot, the oligonucleotide route stops being the backup plan and becomes the main plan. It is also the slowest, which is exactly why it can't wait for the screen to finish.
This is the single most important expectation-setting passage in this primer. Knowing the base rate in advance is what lets you hear a negative screen result as information rather than as catastrophe.
Hit triage — the filters that matter before anyone gets excited:
- Is it real? Dose-dependent, reproducible, and confirmed by an orthogonal method.
- Is it specific, or just toxic? Many "hits" simply make cells sick, which changes every measurement. This is the most common false positive in the field. A proper screen always runs a parallel cytotoxicity/viability assay.
- Does it reach the brain? A drug that can't cross the blood-brain barrier cannot work here, no matter how good the dish data. Filter early.
- Is it safe in children, chronically? A child would potentially take this for years. An oncology drug with serious toxicity is a very different proposition from a well-tolerated approved medicine.
- Is the effective dose achievable in a person? A compound that works at a concentration unreachable in a human brain is not a candidate.
- What's the mechanism? Not strictly required, but a plausible mechanism greatly strengthens the case and guides dosing.
A caution about the emotional dynamics. The word "hit" sounds like a discovery. In screening it means "a signal worth checking." Most hits do not survive triage — that is the normal, expected, correctly-functioning process, not a failure. Knowing this in advance is genuinely protective: it lets you hear "we found some hits" as encouraging news about the process rather than as news about a treatment.
Key terms:
- Hit — a compound producing the desired signal in the primary screen.
- Lead — a hit that has survived confirmation and triage.
- False positive — an apparent hit that isn't real.
- Cytotoxicity — the compound is simply killing or sickening cells.
- Blood-brain barrier (BBB) — the barrier restricting what enters the brain from blood.
- Pharmacokinetics (PK) — what the body does to a drug (absorption, distribution, metabolism, excretion).
- Pharmacodynamics (PD) — what the drug does to the body.
- IC50 / EC50 — the concentration producing half-maximal inhibition/effect.
- Repurposing — using an existing approved drug for a new disease.
- Counter-screen — a parallel assay to rule out trivial explanations such as toxicity.
References:
- Broad Drug Repurposing Hub
- NCATS — National Center for Advancing Translational Sciences
- ReFRAME drug repurposing library