Friday, July 23, 2010

How to Discover


For myself and my future trainees, I am using this blog entry to reflect on the principle components of the path to discovery in bench science. We have a limited lifetime to discover things and make an impact in our careers, so it is worth thinking carefully about the questions and directions one chooses to focus on. Discovery is at its best when it arrives in the form of an unexpected observation, but there is an art to discovery and a correct approach that will position an individual to make unexpected observations. Intertwined with that approach is the right attitude and the art of phrasing an interesting question. I want to comment on both aspects of discovery from my experiences. It is important to note that discovery leads to more problems/questions, as much as it does to solutions. An important part of the scientific process is solving problems and I think of this as separate to discovery. A full scientific research program in a lab balances a spectrum of projects focused on discovery, problem solving, and the generation of new technology within a focused area of expertise. Research projects within the lab should complement each other to promote synergy, rather than dilute expertise and interest across too many poorly related questions. A central bread-and-butter theme is needed to unite the lab environment, especially at early stages.

I start with an outline of the scientific process and end with some personal perspectives.

The Scientific Process
The scientific process is a series of deductive and inductive steps (see Nisbet, Elder and Miner. Handbook of Statistical Analysis and Data Mining Applications,.2009).

1. Define the problem and central question to be answered
2. Gather existing information about the phenomenon
3. Form one or more hypotheses
4. Collect new experimental data
5. Analyze the information in the new data set
6. Interpret the results.
7. Synthesize conclusions, based on the old data, new data, and intuition.
8. Form new hypotheses for further testing
9. Do it again.

Before beginning a major new body of work, it is useful to reflect on aspects of this process in the design of your project. Consider the following checklist:

The Experimenter's CheckList

1. What is the question you asking? Refine this question so it is clear and simple.
2. What have others found with regard to this question? On a scale of from 1 to 10, what are the opportunities for discovery? Is this a crowded, old field?
3. How can you improve your question to make it more interesting? To ask something that has not been asked previously?
4. How does your question fit into a story or conceptual framework? What stories could you tell? Scientific publishing is about the story.  This is simply how human beings understand information.
5. What is interesting or important about your question? Where could it take you in your dreams? Rank it on a scale of 1 to 10.
6. Can you optimize your question to maximize your chances of discovery?
7. What are the different approaches you could take to address the question? Balance Risk vs Reward in the decision for the approach.
8. What controls will you need to interpret your experimental results?
9. How will you analyze your results?! What statistical approaches will you use. Think of this before, so you don't forget controls that prevent interpretation of the data.
10. What preliminary tests could you do to optimize technical aspects of your approach to improve the quality of your data? (Always take the time to optimize so that your results will be clear and interpretable....do not base conclusions on weak results!)
11. Write out each of the steps in your experiment, so you don't need to think while you are doing it!
12. What will contribute to variability in your experiment? How can you control that variability?
13. What will the final figure look like that will present your results? What are the possible comparisons and different ways of looking at the data you will get?
14. What degree of effect do you expect in your study? Do you expect a lot of variability? How many replicates will you need to overcome that variability and detect a real effect?
15. Do you have all the reagents and/or samples (animals) you need for your experiment? Figure this out ahead of time.
16. Plan everything out in day planner schedule to determine how long it will take and when certain milestones will be achieved.
17.  Do not under power your study.  Plan for an appropriately high n, so you can make solid conclusions at the end.
18. Order everything you need and FINISH your experiment! A well conceived experiment is always worth finishing to the END! (always finish your experiments even if you loose heart half way through)

Points to Consider

1. When you think about your experiment from a technical perspective, think about efficiency. At each step you capture a certain percentage of the effect and with each step you will introduce a certain amount of noise. How do you optimize your signal to noise? How big of an effect are you seeking?
2. Make sure your initial results are truly rock SOLID. These are the foundations of your project and if they are flimsy, you will reach a stage where you will feel pressure to bend your data to fit those early observations so you can get that paper together under extreme stress in 3-4 years time.    The pressure is coming....make sure the foundations of your study...your first experiments....are rock solid, so you are set on the right path from the beginning. BE PATIENT if your observations are ambiguous and keep working to NAIL THAT RESULT!
3. Thinking about the final figure that will represent your data is essential at the very beginning, so you know where you are headed. Design your figures mentally before and during an experiment. (I constantly pencil out figures to think about results)
4. FINISH YOUR EXPERIMENT AND IMMEDIATELY MAKE A PUBLICATION QUALITY FIGURE. This ensures you are making progress and on top of your data, even if the results don't make sense at that time. They may make perfect sense years later!

FInal Thoughts:

The Power of Characterizing Biological Systems
Science is a balance of risk and reward. The existing funding system forces you to balance your risk carefully. You must show productivity and at the same time push into new frontiers. Though recipes for discovery are limiting in themselves, I tend to favor characterization projects to get going. Simple approaches that involve just "looking and learning", like observing a behavior pattern or staining for several marker proteins to visualize the organization of neural circuit, can expose lots of new questions. Mouse genetic approaches to characterizing a system can also give great and elegant insights as well, but carry more risk and time investment. An unbiased characterization and careful systematic analysis of the results will foster ideas. Look for new technical approaches to characterizing the system associated with your question. A new angle can change everything!! Always think of new technical approaches to ask questions in ways that could not be done before, you will always learn something. Characterization projects lead to papers and useful knowledge with relatively little risk and they set you up for discovery. Everyone should have a component of their research that simply involves looking and describing.

How to look:
1. Make an unbiased list of the features you are seeing in your data. For example, if you are looking at the expression pattern of several genes of interest, where are they expressed? what types of cells? Where are the cells? What do we know about their functions?
2. Design some specific questions (or hypotheses) from your initial observations. Make a list of several different ideas and questions and use your gut to judge the best place to start.
3. Think of methods to rigorously quantitate and statistically analyze your characterization of the system. For example, quantitate cell numbers, measure dendritic projection patterns, monitor feeding patterns, etc.


Balance Characterization Projects with Innovation Projects
Characterization projects generate new knowledge and I think it is important to distinguish knowledge generation from innovation. Innovation involves solving a problem for the first time or in some novel manner, or following a crazy idea to see where it might lead. Innovation is about following your gut and trying a high risk, high reward idea. You must be comfortable and accept regular failure in the road to innovation and that is why it is important to balance your research program by having two classes of projects. In the optimal circumstances, characterization projects and innovation projects complement each other.


Build A Discovery Niche
Big discoveries are made by:
(1) Doing something others cannot do, because they don't have access to the knowledge, resources and tools necessary - aka. build new tools/resources that only you have and learn new fields whenever possible.
(2) Doing something others won't do, because it is very difficult - no replacement for hardwork
(3) Fortunate insight or chance discovery that is capitalized on (the prepared mind!) - pay attention, do carefully controlled experiments, think deeply about your results
(4) Taking risks. You must take some risks in your career and recognize that the path you start down is rarely headed where you think it is...hold weak opinions.
(5) Reading and talking. You must learn broadly in order to understand the impact of your results and observations and connections to other fields. Something that seems mundane might be huge when cast in the right light.

Think Differently
(1) Characterization projects are powerful ways to develop hypotheses and break new ground, but you must use them strategically. Think differently. Look for untracked territory and ways to bring together different fields.
(2) If you are uncomfortable and unsure of where your work is leading, but you find it very interesting....that is normal and that is life on the front of innovation. Just keep asking good questions.
(3) The genius is in the details. Think carefully about the what, why, where and when details of your observations.

MOST IMPORTANTLY
JUST TRY! JUST KEEP TRYING! Never fear failure or a new technique. Dust yourself off and try again. Science is mostly about just trying....you often won't know if you are on to a good thing until late in the game.


Wednesday, October 28, 2009

Programming :: Mac to Unix :: Excel Exports have New Line Problems for Perl on OSX

Why won't my Perl script parse the CSV or tab-delimited table I just exported from Excel?
Note that when exporting a tab-delimited or CSV file from Excel in MacOSX you may find you can't get your Perl scripts to work on the text file. This is likely due to problems with the newline characters. Excel tab delimited files are exported with carriage returns (\r) rather than the unix linefeeds (\n) at the end of each line.  This can cause a lot of frustration and here are some solutions:
1.  Download the software Tex_Edit Plus.  Open the tab delimited file in Tex_Edit. Go to Tools; Quick Cleanup; Mac to Unix conversion.  Save the file and now run your script on this converted file.
2.  Alternatively, there should be a mac2unix tool in your usr/bin/ that could work in your perl script if you have it.
3.  With Perl you can do a mac to unix conversion using the regular expression $line =~ s/\r/\n/g, but if you are reading in your file through a while loop it is already screwed...so it is easier to do the conversion from the Unix command line by typing the following at the command line:
perl -p -e 's/\r/\n/g' (less than sign)excelexport_infile.txt (greater than sign)unixconverted_outfile.txt
* this regular expression globally (g) replaces (s) returns (\r) with newlines (\n).  Note that html requires that I write out (less than sign) and (greater than sign), but I mean for you to use the actual signs here.

These will switch the newline characters to Unix (\n) and your perl script should now work on the unixconverted_outfile.txt.

Wednesday, April 29, 2009

Introduction to the Kinship Theory


The leading theoretical explanation for the evolution of genomic imprinting is the Kinship Theory, which was proposed by David Haig here at Harvard in 1989. The Kinship theory is founded on work by Bob Trivers and Bill Hamilton, who first introduced the highly influential concepts of ‘parent-offspring conflict’ and ‘inclusive fitness’ into evolutionary biology, respectively. Here I post another fine excerpt from Brady Weissbourd’s undergraduate thesis that gives a nice introduction to the Kinship Theory. Thanks again Brady!!

An Introduction to the Kinship Theory of Genomic Imprinting

Haig’s Kinship Theory requires an understanding of modern evolutionary theory, which at the level of the gene has yielded the surprising finding that, within an individual, genes can be in conflict with each other (Haig, 1989; Haig, 2000). Richard Dawkins (1976) argues that evolution should be viewed from the level of the gene: that individuals are vessels created by genes to serve the agenda of the genes. That agenda is propagation. Indeed, from the original “replicators” to complex genomes, throughout evolution these gene-carrying machines have become increasingly complex and efficient at conveying a gene – or set of genes – to the next generation. In this sense, each gene acts in its own best interest to increase its chances of making it to the next round (Dawkins, 1976).

Despite the language of intent attributed to these genes, it is the blind power of natural selection that causes “selfish” genes to propagate in a population. For a gene to be positively selected for and therefore be successful a population, it needs to confer a benefit to itself and not necessarily to the individual, though the interests of the individual and of the genes tend to align. There are numerous of examples of “selfish genes” that survive at the expense of the individual in both plants and animals. For example, so-called segregation distortion genes (also called meiotic drive) increase their representation in the next generation often at a cost to the individual, such as lower sperm viability (Taylor and Ingvarsson, 2003).

Gene selection may also explain complex cooperative, and even so-called “altruistic” behavior. A gene that promotes the fitness of other individuals who have high probabilities of carrying that gene will also proliferate in the population, so long as the cost to self is less than the benefit to the relative, weighted by the degree of relatedness (rB>C) (Trivers, 1971; Trivers, 1974). That is, a gene acting “selfishly” can incur a cost to its reproductive opportunities so long as it promotes the passing on of that gene to the next generation via relatives. The higher the probability that a relative carries the gene, the more likely social and cooperative behaviors are to be enriched at that locus (Burt and Trivers, 2006).

However, as David Haig said, “my mothers kin are not my father’s kin” (Haig, 1997). Thus, just as relatedness coefficients can promote “altruism”, intragenomic conflict can arise in populations with asymmetric relatedness between the genes of mothers and fathers in an offspring (Burt and Trivers, 2006; Haig, 2000). Trivers’ understanding of relatedness coefficients and his equation for kin selection (rB>C) in asymmetrically related kin groups yields a framework for understanding the evolution of intragenomic conflict, and thus, genomic imprinting (Haig and Westoby, 1989/91; Haig and Trivers, 1995; Trivers, 1971; Trivers, 1974). However, asymmetric relatedness between maternal and paternal lineages exists outside of the taxa that display genomic imprinting, which highlights the special circumstances necessary for imprinting to evolve.

The fundamental origin of conflict that results in imprinting is that the father is not related to the mother (assuming the absence of inbreeding). In therian mating systems, particularly in those mammals with an invasive placenta, the process of carrying, birthing, and nursing young bears heavy costs to the mother. For each unit the mother invests in a current offspring (e.g. in the form of time or energy), there is a significant opportunity cost to the mother’s future reproductive options. For example, this opportunity cost can manifest as reduced energy available for the current or next litter, or a lower chance of survival until the next breeding season. Mothers will therefore favor some optimal investment in the current generation that leaves her ample energy to survive and invest in future litters (Haig, 2000; Haig, 2002; Haig, 2003).

The optimal maternal investment from the point of view of the father is quite different. It is often the case that females will mate with multiple males, both within and across breeding seasons. Indeed, litters of pups (or even human twins (Ambach et al, 2000)) may have multiple fathers. Therefore, given that fathers and mothers are not genetically related, the future reproductive interests of the mother are of little to no consequence to the father after the birth of their offspring. The maternal optimum takes into account the survival of each member of the current generation as well as future reproductive fitness, while fathers care only for offspring in the current generation that carry a paternally derived allele. This results in fathers preferring a higher level of investment in the current generation, specifically in any one offspring carrying a paternal allele (Haig, 2000; Haig, 2002; Haig and Wilkins, 2003).

It is then expected that paternally inherited (padumnal) genes will upregulate offspring behaviors and physiological processes that tax the mother beyond her optimum, while maternally inherited (madumnal) genes will evolve to counteract this strategy in an evolutionary “arms race”. An offspring’s madumnal genes are related to the offspring by 1, and related to the mother by 1 (retrospectively, they are maternally inherited), and therefore are selected to support maternal interests. Similarly, padumnal genes are related to offspring and father by 1, which creates an immediate conflict in interest between madumnal and padumnal genes in their interactions with kin. As these madumnal and padumnal genes come together to form a temporary vessel, they have a common interest in the reproductive potential of the organism they cohabitate. However, if a gene carries an imprint of which lineage it represents, that gene may be selected to serve the marginally different interests of the paternal versus maternal faction (Haig, 2000; Haig and Wharton, 2003).

A gene will therefore become imprinted when madumnal and padumnal genes have different optima that serve their different interests. The evolutionarily stable strategy is then to silence the gene that prefers the lower value, and express the other at its optimum. As genes in the offspring are selected to demand more resources, the corresponding maternal genes are silenced. Conversely, maternal genes suppressing resource acquisition will approach maternal optimums, accompanied by paternal silencing (Haig and Wharton, 2003).

This inherent difference in relatedness between maternally and paternally inherited genes may also apply to sibling interactions. Madumnal genes in offspring are confident of their relatedness to siblings, and mothers are confident of their equal relatedness to each pup, so selection for cooperative behaviors and equal provisioning of pups is predicted. Conversely, paternity certainty is low, and padumnal genes that enhance the intake of maternal resources, even at a cost to siblings, are predicted to spread through the population. It is also generally the case that mammals live in systems where matrilineal groups tend to stay together, with male dispersal (Haig, 1997; Haig, 2001). Therefore, conflict between “cooperative” behaviors and “selfish” behaviors will be represented by maternal and paternal lineages respectively. It is then expected that the degree of conflict will increase with increased promiscuity in the population, and decrease with increased monogamy (Burt and Trivers, 2006; Haig, 2000; Haig and Wharton, 2003).

This theory comprehensively explains the evolution of imprinting in placental mammals and provides a compelling framework for the study of the physiological and behavioral effects of imprinted genes. Further, the formulation of kinship theory that takes into account how imprinting can arise from the asymmetries of relatedness inherent in mammalian social groups – not just in the relationship between mothers and their offspring – suggests that imprinted genes likely have evolved to control complex social interactions and behaviors (Haig, 2000).
Evidence from known imprinted phenomena strongly supports the kinship theory of imprinting. The preliminary studies of failed development in parthenotes demonstrate that in androgenotes, there is enhanced development of extraembryonic tissue – those responsible for the acquisition of resources from the mother – with impaired growth of the embryo itself. Conversely, gynogenotes have severe deficits in extraembryonic tissue, with relatively normal embryonic growth (Surani et al, 1984). Further, insulin-like growth factor 2 (IGF2) is known to promote fetal growth while IGF2 receptor (IGF2R) reduces concentrations of IGF2, reducing growth. As predicted, IGF2 is a paternally expressed gene (PEG) and IGF2R is a maternally expressed gene (MEG). Other PEGs such as Kcnq1ot1 have been shown to increase growth while a number of MEGs, such as cdkn1c, H19, PHLDA2 are known to suppress it (Haig, 2004).

Brandon Weissbourd, Harvard Undergraduate Honors Thesis, 2009

What is Genomic Imprinting?


Genomic imprinting is one of the most provocative and exciting fields of research falling under the umbrella of “epigenetics”. Imprinting is thought to be a rare, but extraordinarily important mode of gene regulation in the genome and is the primary focus of my research. For this post, I am putting up an excerpt from Brady Weissbourd’s undergraduate thesis. Brady is an outstanding undergraduate here at Harvard who has worked with me for the past year and presents a concise and interesting introduction to the imprinting field in his thesis. I hope you find it informative. Thanks Brady!!

The Discovery of Imprinting

A set of experiments in the 1980’s challenged traditional views of genetics, rewriting Mendelian paradigms and launching the field of genomic imprinting. Despite the expectation that genes are ignorant of their inheritance, it was shown by McGrath and Solter (1984), as well as Surani et al (1984), that both maternally and paternally inherited alleles are required for normal development. Surani et al (1984) demonstrated this phenomenon by making heterozygous and uniparental embryos via nuclear transplantation by adding either a maternal or paternal pronucleus to haploid parthenogenic (unfertilized) embryos. If a paternal pronucleus was added, embryos were able to develop to term. However, eggs with two maternal pronuclei showed embryonic growth, but failed to develop extraembryonic tissues and trophoblast, and survived no longer than the 25-somite stage, roughly ten days after conception (Surani et al, 1984).



This demonstrated that a paternally inherited genome is essential for the development of extraembryonic tissue. Further, embryos created with two paternal pronuclei showed that maternally derived genes are essential for the development of the embryo: androgenetic embryos developed an almost normal trophoblast, but the embryo proper was severely retarded (McGrath and Solter, 1984; Solter, 1988; Surani et al, 1984). Interestingly, human androgenetic cells can develop into hydatidiform moles, characterized by extensive overgrowth of extraembryonic tissues, but generally devoid of an embryo (Bagshawe and Lawler, 1982).

Further studies by Cattanach (1986) and Cattanach and Kirk (1985) of uniparental disomy in mice demonstrated that different phenotypes result from the loss of the maternally versus paternally inherited allele at a single locus. For example, Cattanach and Kirk (1985) showed that within the same litter, a paternal duplication on chromosome 11 lead to large pups, while a maternal duplication at the same locus lead to small pups. Large and small pups for paternal and maternal duplications respectively were also observed for the distal region of chromosome 2. Further, disomies on chromosomes 2, 6, 7, and 8 were shown to result in pre-natal lethality (Cattanach, 1986). Interestingly, maternal disomy 6 is lethal, while paternal disomy 6 is normal (Cattanach and Kirk, 1985; Searle and Beechey, 1985).

These mysterious findings have opened new avenues of research and generated major breakthroughs in the study of genetic inheritance. Indeed, in the last twenty years it has been shown that a subset of genes carry “imprints” of their parent of origin, and are differentially expressed depending on whether they were inherited from the mother or the father (Wilkins and Haig, 2003). These imprints are of epigenetic origin, meaning that the observed heritable differences do not reflect differences in DNA sequence but rather a specific regulatory mechanism, most often in the form of histone modification, DNA methylation, and binding of non-coding RNAs (Tang and Ho, 2007).

Imprinting is a phenomenon hitherto observed only in plants, marsupials and placental mammals (therians) (Haig and Trivers, 1995). This taxonomic distribution suggests that it specifically evolved in tandem with changing parent-offspring relationships (Reik, 2001). The exception to this observation is found in viviparous scale insects and sciarid flies, in which only the maternal chromosomes are present in sperm, as the maternal genome eliminates the paternal contribution during spermatogenesis (Haig, 1992; Haig, 1993; Haig and Trivers, 1995).

The discovery of imprinting has provoked exciting new experimental research and theoretical considerations, providing profound insight into the elegance of evolution and the intricacies of mammalian cognition and behavior.

Brandon Weissbourd, Harvard Honors Thesis, 2009

Thursday, December 25, 2008

The Male Brain and The Female Brain Get Confused


For some, the term “neural conflict” might arouse images of severe headache pain or perhaps situational conundrums such as whether or not to have just one more piece of chocolate cake. However, I define neural conflicts in terms of brain circuits that function in opposition to each other, such that, in many situations, both systems cannot be actively commanding an individual’s behavior at the same time. For example, you cannot be awake and asleep at the same time, nor can you be hungry and satiated, or both suffering miserably and reviling in hedonistic pleasure. In this entry, I highlight a bizarre form of neural conflict….a conflict between the male side of the brain and the female side of the brain. Differences in the mental functioning of men versus women are touchy subjects, but remain a vigorous area of research. Over the years subtle differences have been suggested. Many of us are familiar with concepts that emerged from psychology regarding the male bias towards spatial and logical/mathematical reasoning versus the female bias toward emotional and verbal skills. Indeed, these famous studies laid the foundation for a theory called the “male brain hypothesis” for the underlying causes of autism, a spectrum of neurological disorders in which social behaviors are impaired, but spatial and mathematical skills are normal or sometimes remarkable. Simon Baron-Cohen suggested that autism spectrum disorders, which are more common in males and characterized by impaired social behaviors, might represent an overtly male brain. In addition, subtle differences in brain size have been observed, with men having a slightly higher percentage of white matter (myelinated areas, see image above) and women a slightly higher percentage of grey matter (see image above, Allen et al., 2003). Further, some regions, such as a structure in the preoptic area of the hypothalamus, appear to be larger in males than females. However, after decades of work we know little about why one person thumps their masculine chest, while another gets in touch with their feminine side. There is no penis like structure hanging off the front of the male brain that explains it all.

After years of research looking for differences between males and females, who would have guessed that the circuits for male and female behavior are both residing happily together in the same brain. That is what is suggested by Catherine Dulac’s study in Nature (Kimhi et al. 2008). They present evidence in mice that in a given individual both male and female circuits exist simultaneously and propose that switches at the level of neuronal circuits regulate incoming sensory information and channel it to the male circuits, if you are a male, and to the female circuits, if you are a female. Their evidence is striking (see the movie below). They used genetic manipulations to silence the sensory "switch" in the brains of female mice. Specifically, the mice were engineered so that an ion channel in their vomeronasal system is deleted so that this organ does not work. The vomeronasal system does not exist in humans, but acts sort of like the sense of smell for rodents and many other animals, primarily to detect pheromones and other social cues. As shown in the attached movie, the female mice with the silent switch and impaired ability to interpret phermones begin to behave like male mice. Remarkably, they mount and fully copulate with other females and males, even though they don’t have a penis.

This is a striking example of neuronal circuits existing in functional conflict with each other. The inappropriate activation of male circuitry in the female brain, resulting from a disruption in the processing of sensory information (pheromones), leads to an inappropriate male-typical pattern of behavior. Identifying the identity and organization of male-typical and female-typical circuits and how they interface with circuits that communicate sensory information from the environment to the brain will be a critical future direction.