logoalt Hacker News

stoppingtoday at 12:53 AM0 repliesview on HN

I think you have it completely backwards. A smaller genome makes small tweaks much more likely to produce measurable effects, which means your data quality is much better. A larger genome with more gene interaction is more likely to produce non-measurable or non-subtle effects with perturbations of a single parameter. If a positive effect can only be measured by simultaneously mutating multiple inter-dependent genes, and you don't have a data point with those mutations, the effect can never be predicted. The model may even predict the exact opposite: if you have a bunch of data points showing that varying one gene at a time causes the organism to die, the model will predict that the sum of those mutations also causes death. With larger genomes you need exponentially more data points to discover subtle effects.