- Research Article
5
- 10.2217/pme.13.22
The 1000 Genomes Project: paving the way for personalized genomic medicine.
- Jun 01, 2013
- Personalized Medicine
- Ian B Gibson + 2 more +2
The 1000 Genomes Project: paving the way for personalized genomic medicine.
Abstract Technical advances in molecular genetics that were being developed in the 1970s led to the initiation of the Human Genome Project (HGP), a long-term international research effort to completely map and sequence the human genome. In 1988, U.S. Congress appropriated funds to the Department of Energy (DOE) and the National Institutes of Health (NIH) to begin the planning stages of this vast endeavor (Collins, 1999). The HGP objectives included construction of a detailed genetic and physical map of the human genome, determination of the complete nucleotide sequence of human DNA, and localization of the currently estimated 30,000 genes within the human genome. The completed HGP therefore comprises a resource of detailed information about the structure, organization, and function of human DNA. Similar analyses of the genomes of other model organisms used extensively in research laboratories are also being performed. This information will change the way medical services are provided in the future, including the treatment of individuals with oral-facial clefting disorders. In addition to promoting the development of improved technologies for biomedical research and genomic analysis, project goals included training scientists to utilize tools and resources developed through the HGP to pursue biological studies that would ultimately improve human health. Specific additional benefits include an enhanced understanding of the genetic contributions to human disease and the development of rational strategies for minimizing or preventing disease phenotypes. Sequencing the human genome greatly improves our understanding of the genetic and molecular basis of disease mechanisms and suggests new therapeutic regimens for disease prevention or, at the very least, the targeting of specific treatment modalities to the individual patient on the basis of their genotype.
The 1000 Genomes Project: paving the way for personalized genomic medicine.
The 1000 Genomes Project: paving the way for personalized genomic medicine.
Completion of human Chromosome 21, the Human Genome Project, and Steps towards Understanding Ourselves through Comparative Genomics
Chromosome 21 is the smallest human chromosome and represents a model for physical mapping of the entire human genome. Three copies of this autosome cause Down syndrome, the most frequent genetic disorder associated with significant mental retardation. Five years ago, in the framework of the international human genome project, and as an integrated part of the chromosome 21 community, a consortium, of academic groups from Japan and Germany was formed to map and sequence this chromosome. Several whole-genome and chromosome-specific DNA libraries were constructed for mapping purposes. Using a combination of nested-deletion [1] and shotgun sequencing approaches, our center determined and analyzed over 17 million bases of high-quality data. In total, over 33.5 million basepairs of DNA, distributed over four contigs, were sequenced with very high accuracy [2]. The largest of these contigs is nearly 25.5Mb long. Only three small clone gaps remain, which together comprise approximately 100kb. Thus, we achieved a coverage of 99.7% of 21q. In addition, 211,116 bp from the short arm were also sequenced. Analysis of the chromosome revealed far fewer genes than previously estimated. Instead of 800-1000 expected genes, we found only 225 genes, of which 127 are known genes and 98 are predicted genes. In addition, 59 pseudogenes were identified. The completion of the 21q sequence provides a unique resource for understanding the molecular pathophysiology of Down syndrome, as well as all other monogenic and complex disorders that map to this chromosome, including Alzheimer’s disease, leukemia, autoimmune disease, epilepsy and manic-depressive psychosis. It also stands as a structural framework from which the complete molecular architecture of the chromosome can be determined. With the announcement of the first “working draft” of the human genome by the international human genome sequencing consortium this past June, there has been a great interest in the total number of genes in the human genome[3]. Some of the current estimates range from as low as 34,000 genes to as high as 140,000 genes. The low gene-density of chromosome 21 was quite unexpected. Most striking is a 7-Mb region near the region near the centromere that contains only one gene. This region is much larger than the whole genome of some species, such as Escherichia coli, yet evolutionary processes permitted the existence of such a gene-poor SNA segment. This finding leads us to propose that there are additional large gene-less regions in other human/mammalian chromosomes. By combining the gene numbers of chromosomes 21 and 22 [4] (770genes; 2-3% of human genome) and assuming that together they represent the average gene content of the human genome, we estimate that the total number of human genes may be close to 40,000. As part of the international human genome sequencing project, we are also sequencing parts of chromosomes 11 and 18. Our high-throughput sequence production line generates about 10 Mb per day. In order to handle such a high volume of output, we have developed a system for automated data assembly, annotation and release. All the data is released through our website and through the DNA Database of Japan immediately according to the policies established by the international consortium. In addition to chromosomes 11, 18 and 21, we are also interested in the structure of the entire human genome. In collaboration with the University of Tokyo, we have developed a database called HGREP [5] (Human Genome Reconstruction Project) that contains working draft and finished sequences which cover more than 85% of the human genomic sequences. Sequence entries are aligned along the chromosomes based on sequence similarities to STS markers, BAC-end and other entry sequences. Further, biological features, such as genes, gene functions, repeats and CpG islands, are fully annotated on these entries. To identify evolutionally conserved and biologically important information in the genome, we are taking the strategy of comparative genomic sequence analysis between human and other related genomes such as those of rodents and primates. We are currently sequencing several regions of the mouse, including some regions that are counterpart to human chromosome 21, and are preparing genomic resources for sequencing of the chimpanzee. Through comparative methods, we are also trying to understand telomeric regions, structures at the end of chromosomes that play an important role in several biological functions and have been associated with human diseases like an oncogenesis and a cellular aging. The structure and function of all genes and their regulatory regions can only be fully understood by looking at how they interact under different circumstances and in several different backgrounds. Determining which genes and elements are human-specific will lead to a deeper understanding of who we are and will advance medical sciences drastically.
Read moreA budget out of balance.
T he new administration's science budget, sketchily outlined in a request to Congress, brought March in like a lion for the National Institutes of Health (NIH). Although that is likely to please our biomedical readers, the budget will disappoint almost everyone else. But from the unfortunates, the silence has been deafening, at least so far. The Battle of the Budget has scarcely been joined, after all, and the advocates for more of this and that are treading carefully. They have abundant reasons for their caution: Although it is a brand-new White House, it has already proven to be hard to move and willing to punish. Thus, the tough questions about the care and feeding of science have, for the most part, not yet been asked. The National Science Foundation's (NSF's) Rita Colwell tactfully eschewed complaints about her agency's miserly less-than-cost-of-living increase, and other heads of agencies have been equally silent about their low budget marks. Even the lobbyists are keeping their powder dry for a while. The happiest camper in the entire government is surely the secretary of Health and Human Services, whose 14% boost for NIH presumably got him high-fived all over the NIH campus during his triumphal budget-day tour. We'll get to the unhappiest campers later. In the meanwhile, since Science (the journal, that is) receives nothing in this budget and therefore has nothing to lose, it seems safe for us to evaluate this proposal in terms of its balance. Does it really make sense for some pieces of the enterprise to be treated very well indeed and others to be held back or cut? There are good reasons for thinking it doesn't. In the first place, an increasing proportion of the important problems in science are interdisciplinary in character. At Science , we have published contributions to nanotechnology that come from disciplines as diverse as chemistry, materials science, and electrical engineering. The climate sciences, on which we will depend in formulating international policies, draw from paleontology, oceanography, and atmospheric chemistry. The dramatic scientific gains that will flow from the sequencing of the human genome will be harvested not only by molecular biologists but also by specialists in bioinformatics, trained in such disciplines as mathematics and computer science. Nurturing fields such as these requires a balanced portfolio. And the balance has to come from a thoughtful cross-cut of the entire federal science budget, by those who know the enterprise well. It is hardly surprising that the balance is lacking here, because the offices that could supply such oversight (most notably, that of the director of the Office of Science and Technology Policy, the president's Science Adviser) are dark. It is well past time for this administration to turn the lights on. As things stand, though, NSF will, unless protests avail, limp along on an increase of 1.2%: an actual reduction in constant dollars once the nonresearch increases are deducted. The Department of Energy (DOE) actually loses 3%. These marks portend difficulties for the physical sciences, but that is not all. Less attended to, but also important, is the prospective impact on the plant sciences. Those disciplines, heavily dependent on both NSF and DOE as well as the flat-lined Department of Agriculture, will provide the science essential to feeding a population due to reach 10 billion by mid-century. That's a worry for the rest of the world, where the needs will strike hardest, and also for U.S. agriculture, which will be heavily depended on to meet them. So where are the unhappiest campers of all? Look for them in the U.S. Geological Survey (USGS), originally targeted for a monster cut of 22%, subsequently reduced to “only” 11%. This agency has supplied most of the government research that has guided fossil fuel exploration for decades. The year of a power production crisis seems an odd time for a cutback. The most compelling irony, though, has to do with another administration plan. USGS, it turns out, has supplied the framework geology that will enable—guess what?—drilling for oil in the Arctic National Wildlife Refuge! Environmentalists may wonder: Could USGS meet its reduction target by going back and undoing that work? Afraid not.
Read moreThe new biology enters the generalist pediatrician's office: lessons from the Human Genome Project.
1. Edward R.B. McCabe, MD, PhD* 1. 2. *Physician-in-Chief, Mattel Children’s Hospital at UCLA; Professor and Executive Chair, Department of Pediatrics, UCLA School of Medicine, Los Angeles, CA. Birth defects are the leading cause of infant mortality in the United States, representing more than 20% of all infant deaths. This infant mortality rate from birth defects exceeds that from sudden infant death syndrome, low-birthweight/short gestation, respiratory distress syndrome, and maternal complications. In addition, birth defects and genetic diseases represent major sources of morbidity for those who survive. As our ability increases to care effectively for those who have infectious diseases and other acute illnesses, individuals who have chronic illnesses due to genetic etiologies represent an increasing proportion of patients seen in the general pediatrician”s office. The Human Genome Project was initiated on October 1, 1990, and has a projected funding period of 15 years. The goal is to sequence the entire human genome, representing three billion base pairs that contain the coding sequences for approximately 75,000 genes. During the latter half of this century, investigations into the genetics of disease gathered increasing momentum. In addition to fundamental investigations into human genetics, technologic tools were developed that permitted large-scale genomic sequencing. These tools included the polymerase chain reaction (PCR), which permits amplification of hundreds of thousands or even millions of copies of DNA and requires only limited sequence data for its success; automated DNA sequencing, which allows increased sequence processing and decreased cost compared with manual methods; and improved information systems, which permit sophisticated analysis and assembly of the three billion base pairs of DNA in the human genome. Thus, the Human Genome Project represents the current chapter in our understanding, but it is neither the first nor the final chapter in this story. Once we know the sequences of all of the human genes, we must learn their functional roles in human development and disease pathogenesis. The Human Genome Project has been referred to as the “moon shot …
Read moreContemplating the End of the Beginning
On February 12, 2001, an unprecedented collection of papers describing the initial sequencing and analysis of the human genome was published in Nature and Science (International Human Genome Sequencing Consortium 2001; Venter et al. 2001). Although much celebration and press attention had surrounded the earlier announcement in June 2000 of the coverage of the vast majority of the human genome sequence in working draft form, the publications in February 2001 carried with them the kind of satisfying scientific significance that laborers in the genome fields had longed for—the full description of the methods used to determine the letters of over 90% of the human instruction book, and a host of surprising revelations from the analysis of its contents. This brief essay represents a personal reflection on how we got here, and where we are going. The achievement of these landmarks, coming years ahead of the original schedule, was only possible because of the advances begun early in the preceding decade, reflecting the polyphonic set of interconnected goals that the planners of the Human Genome Project (HGP) wisely included as part of the original master plan. Science traditionally operates by the process of researchers standing on the shoulders of those who came before, and that has certainly been true for the HGP. Building detailed genetic and physical maps, developing better, cheaper, and faster technologies for handling DNA, and mapping and sequencing the more modest-sized genomes of model organisms were all critical stepping stones on the path to initiating the large-scale sequencing of the human genome. Pilot efforts to sequence the human genome began in the mid-1990s. When the International Human Genome Sequencing Consortium met for the first time in Bermuda in 1996, there was a sense of excitement, but the magnitude of the task at hand was sobering– throughput was too low, costs were too high, technology was still immature. Despite that anxiety, the assembled scientific leaders from several countries at that meeting endorsed the importance of high quality sequence, and made one of the most crucial decisions of the genome era–immediate data release. Led by John Sulston and BobWaterston, whohad adopted this same policy for the sequence of C. elegans, the assembled sequencing center directors unanimously adopted a statement that all assembled contigs greater than 1 or 2 kb would be placed in public databases within 24 hours. The argument was simple: The sequence would only benefit the public fully if it could be understood, and that required making it immediately available so that all the creative minds of the planet could work on it. The establishment of this principle was one of the defining moments of the HGP. Over the next three years the rate-limiting steps of large-scale sequencing began to yield to creative innovations. The genome centers implemented major improvements in library production, template preparation, and laboratory information management, so that less and less human intervention was required in the main production pipelines. The advent of capillary sequencing machines from Amersham and ABD provided a much-needed boost in efficiency. Much has also been made of the appearance of a commercial entity on the scene in May 1998 (Celera Genomics) as an additional nudge to the HGP. Whereas it is fair to say that the resulting sense of competition provided an additional incentive to the genome centers, it would be misguided to say that the HGP was previously operating in a relaxed fashion, or that the significant advances in throughput would otherwise not have happened. After all, most of those advances were born of previous accomplishments of the HGP itself. From my perspective, a major turning point occurred in Houston in February 1999. The largest NIHfunded centers (at the Whitehead Institute, Washington University, and Baylor College of Medicine) had just undergone rigorous peer review of their proposals to scale up sequencing throughput and were about to receive a significant increase in funding. The Sanger Centre in Hinxton (UK) and the Joint Genome Institute (Walnut Creek, CA) of the Department of Energy were also scaling up production rapidly. An experiment carried out the preceding summer had documented the high degree of utility of a “working draft” human genome sequence; the half-dozen labs that compared draft and finished sequences found that the draft could answer most of the scientific queries they posed (although it was harder to work with), and suggested that the majority of the HGP’s efforts might well be devoted to obtaining working-draft coverage of the genome as quickly as possible, as long as the commitment to finishing was not diminished. Accordingly, the sequencing plans for the NIH and DOE that were E-MAIL: fc23a@nih.gov; FAX (301) 402-0837. Article and publication are at http://www.genome.org/cgi/doi/10.1101/ gr.1898. Commentary
Read moreThe New Genetics and Women
The U.S. Human Genome Project (HGP) is a federally funded effort to produce a detailed genetic and physical map of all human chromosomes. Because of their central role in reproduction and caregiving, women are likely to be affected differently, and more significantly, by the information the HGP generates. It is important to identify inequities that may emerge from gender differences and to consider ways in which they may be avoided, reduced, or overcome. Although this type of analysis is one of the goals of the Ethical, Legal, and Social Issues Program of the National Center for Human Genome Research, few studies have focused explicitly on the impact of the HGP on women. This article describes the potential impact of the "new genetics" on women. Identification of gender differences as they affect both research and clinical practice and the psychosocial, legal, and ethical implications of the HGP should evoke and inform public discussion and policies that may be generated by these issues.
Read moreIntroduction to Molecular Genetics and Genetic Testing for Retinal Dystrophies
The study of the inherited dystrophies represents one of the greatest success stories of modern molecular genetics. This is a story that began with the description of linkage to part of the short arm of the X chromosome (described as early as 1985 for one of the loci for X-linked retinitis pigmentosa) and the identification of the first gene known to be mutated in adRP (encoding the polypeptide rhodopsin) in 1990. The successful international effort to sequence the entire human genome – the Human Genome Project – represented a landmark collaborative project and its completion, which has enabled the entire human genome to be mapped, further accelerated the ability of scientists to understand the underlying molecular basis of human genetic disease. Consequently, at the time of writing over 200 gene loci and 150 genes have been described underlying human monogenic retinal disorders implying a level of complexity that was unsuspected 20 years ago. This has allowed an entirely new understanding of the genetic basis of this group of conditions. The ‘genetic’ conditions referred to in this text, and within this section, will be monogenic, or Mendelian, conditions. However, it is now recognised that many common conditions have a substantial genetic contribution; the recent delineation of polymorphisms in the complement pathway as key genetic contributors to age-related macular degeneration demonstrates that molecular genetic discoveries have by no means been limited to the Mendelian retinal dystrophies.
Read moreHuman genome project Europe
Human genome project Europe
Heredity under the Microscope: Chromosomes and the Study of the Human Genome by Soraya de Chadarevian (review).
Reviewed by: Heredity under the Microscope: Chromosomes and the Study of the Human Genome by Soraya de Chadarevian Miguel García-Sancho Soraya de Chadarevian. Heredity under the Microscope: Chromosomes and the Study of the Human Genome. Chicago: University of Chicago Press, 2020. 272 pp. Ill. $37.50 (978-0-226-68511-3). In Heredity under the Microscope, Soraya de Chadarevian complements the picture of post-World War II biomedicine that she uncovered in her previous book, Designs for Life: Molecular Biology after World War II.1 Unlike in her earlier monograph, here she focuses on chromosomes as spaces of convergence of several lines of inquiry. This focus on research objects rather than disciplinary developments leads her to investigate a variety of practices that life scientists deployed to visualize and [End Page 466] interpret human chromosomes. The diversity of disciplinary backgrounds of these scientists and the multiplicity of uses to which their chromosome practices were put enable the author to expand the historiographical boundaries of both molecular biology and genetics research. Firstly, de Chadarevian disentangles “the study of postwar human heredity from the predominant concern about continuities with eugenic practices” (p. 5). The book does not limit itself to documenting the invention of karyotyping techniques to represent and detect anomalies in human chromosomes; it follows their standardization and spread, including their use as evidence of safe levels of radiation in the workplace (chapter 1) or sex assignment in the Olympic games (chapter 3). The emphasis on these uses, departed as they were from scientific theories about population improvement, displaces eugenics as the dominant interpretative framework of the history of human genetics. Secondly, the book presents chromosomes as showing “historical commonalities” (p. 177) between two approaches that have been traditionally considered as separate: molecular genetics and microscope-assisted chromosome observations. De Chadarevian stresses how these two approaches emerged as a consequence of the key role that nuclear research had played during World War II and the search of peacetime applications. Radioactive isotopes decisively fostered molecular biology, while the investigation of the effects of radiation was a crucial concern of human and medical geneticists using chromosome observations (chapter 1). Both molecular biologists and chromosome researchers also attempted to apply computers to their investigations, the former as an aid to the structural study of proteins and DNA, and the latter to automate the interpretation of karyotype images (chapter 4). However, whereas molecular biologists focused on simpler organisms, chromosome researchers engaged with both human populations and patients in the clinic. Radiation-induced leukemia was their first focus for then rapidly shifting to other conditions (especially those affecting mental abilities and sex identity) and larger-scale epidemiological studies (chapters 2 and 4). The 1960s, a decade that historians have dubbed the golden age of molecular biology,2 was also a time of expansion of chromosome observations. From the mid-1970s onward, molecular biologists turned to human subjects in order to realize the medical promise of the new recombinant DNA techniques. It is at this point when the intersection with chromosome research becomes more visible and de Chadarevian wonders “whose turn” this represents in history, a molecularization of human genetics or a humanization of molecular biology (pp. 171–76). Although subsequent success stories have put the emphasis on molecular biology, chapter 5 shows the importance of chromosome research as a key tool for gene mapping and the emergence of the human genome as an object of study. Even in the so-called post-genomic era, chromosome observations are [End Page 467] essential to clinically interpret the molecular sequence data derived from the Human Genome Project. Overall, the book offers a most welcomed framework to address the coexistence of chromosome observations and molecular biology, and their importance in genomics research. Understanding this is a pressing historiographical problem, given the dominance of molecular biologists in existing narratives and the apparent revival of chromosome research. There is yet still work to do in exploring specific overlaps between the visual and molecular approaches to the genome. A potential candidate is Victor McKusick, a main actor in de Chadarevian’s account and a co-author of the article in which Celera Genomics presented one of the draft sequences of the human...
Read moreThe human genome sequence expedition: views from the "base camp".
The past year has brought unprecedented public attention to biomedical research, with a particularly intense focus on the Human Genome Project and the completion of a first-generation ∼3-billion-basepair human genome sequence. Much of this attention related to the competition between the two parallel, yet separate, efforts of the publicly-funded International Human Genome Sequencing Consortium and the private company Celera Genomics. Despite the apparent rancor between the groups, two celebratory events notably punctuate the past year: the joint media announcement in late June 2000 that both groups had generated a “working draft” sequence of the human genome and the two landmark scientific publications in February 2001 that describe the efforts of each project (International Human Genome Sequencing Consortium 2001; Venter et al. 2001). Numerous grandiose cliches and metaphors have been used to convey the magnitude of these accomplishments and their associated implications for biomedical research and clinical medicine. Here we add one more to this list. Our choice for capturing the essence of contemporary human genome analysis is an analogy to a mountain climbing expedition, one where significant progress has been made to provide a spectacular view of the genetic landscape. But this is not an expedition that is complete, with uncertain— yet exciting—genomic terrain ahead. Indeed, the Human Genome Project is now firmly at the “base camp” of the expedition to elucidate the human genetic blueprint and to begin to understand its content. Nevertheless, this is a milestone of tremendous significance and excitement. Here we outline some of the key lessons learned during the initial analysis of the human genome sequence. We highlight a few of the many remaining questions in understanding the genome’s structure and function, with most answers likely becoming available later in the expedition. Finally, we preview the anticipated climb to the final summit and the ascent to a complete and finished sequence.
Read more19 Emerging Technologies in DNA Sequencing
More than just a mapping and sequencing endeavor, the Human Genome Project (HGP) has altered the mindset and approach to many basic and applied research efforts. Early skepticism and controversy (Koshland 1989; Luria et al. 1989; Roberts 1989b; Fox et al. 1990) were soon laid to rest by well-developed strategies (Roberts 1989a; Collins and Galas 1993; Collins et al. 1998) that led to the successful execution of mankind’s largest biology project. At the core of the HGP was technology development that advanced the pace of sequencing a mammalian-size genome from years to months. Along the way, numerous strategies emerged that hold promise for rapid, efficient, and inexpensive delivery of DNA sequence information. For the HGP, a brute-force approach was adopted for completing the job by coupling the core technologies of Sanger sequencing and fluorescence detection. The completion of the sequencing phase could not have been accomplished without major innovations in recombinant protein engineering, fluorescent dye development, capillary electrophoresis, automation, robotics, informatics, and process management. The result was completion of a high-quality, reference sequence of the human genome in April, 2003 (Collins et al. 2003), marking the 50-year anniversary of the discovery of the double-helix structure. For many outside the genome community, that heroic milestone signaled the end of this international scientific project, but for the rest of us, it only marked the beginning of things to come. The need for sequencing has never been greater than it is today, with applications spanning diverse research sectors including comparative genomics and evolution...
Read moreThe Use of Cytokine Knockouts to Study Host Defense Against Infection
The sequencing of the human genome has revealed the existence of a large number of new members of the various cytokine families and has raised the perennial issue of how cytokines function to benefit human evolution. Usually duplication or redundancy of a gene suggests that it functions to assist the survival of the species. In the case of the proinflammatory cytokines, their role in the survival of the host is less clear since this class of cytokines is clearly implicated in host defense mechanisms against infection as well as the pathogenic processes of several diseases. The interleukin-1 (IL-1) and tumor necrosis factor (TNF) families of proinflammatory cytokines have been studied in great detail for their roles in disease; however, studies on their roles in host defense against infection are limited to a few organisms and models. However, it remains likely that these cytokines contribute to host defense mechanisms using the very same molecular and cellular pathways involved in chronic inflammation, tissue destruction, and tissue remodeling. How does one reconcile the existence of two opposing outcomes for the host—one clearly not in the interest of survival and another vital to the host in order to survive even a minor invasion by a microbe? In each case, identical molecular mechanisms are at work. For example, the processes for emigration of neutrophils into tissues to engulf microorganisms via an increase in endothelial adhesion molecules and chemokine production are the same for the host whether in infected tissue or an inflamed joint. On closer examination, the production and activity of proinflammatory cytokines reveal several levels of regulation, which may explain these seemingly discordant roles for the host. For routine infection or injury, cytokines such as IL-1β and TNF-α are transiently expressed and secreted, their activities are modulated by coproduction of naturally occurring antiinflammatory cytokines, and the levels of production fall rapidly. Thus, IL-1β and TNF-α turn on and off rapidly during infection. In contrast, the production of these same cytokines in chronic inflammatory diseases such as rheumatoid arthritis and inflammatory bowel disease fails to cease. Cytokine production in host defense is regulated, whereas in chronic disease it is dysregulated.
Read moreKnowledge Infusion: In Pursuit of Robustness in Artificial Intelligence
Endowing computers with the ability to apply commonsense knowledge with human- level performance is a primary challenge for computer science, comparable in importance to past great challenges in other fields of science such as the sequencing of the human genome. The right approach to this problem is still under debate. Here we shall discuss and attempt to justify one ap- proach, that of knowledge infusion. This approach is based on the view that the fundamental objective that needs to be achieved is robustness in the following sense: a framework is needed in which a computer system can represent pieces of knowledge about the world, each piece having some un- certainty, and the interactions among the pieces having even more uncertainty, such that the system can nevertheless reason from these pieces so that the uncertainties in its conclusions are at least controlled. In knowledge infusion rules are learned from the world in a principled way so that sub- sequent reasoning using these rules will also be principled, and subject only to errors that can be bounded in terms of the inverse of the effort invested in the learning process.
Read moreCurrent trends in mapping human genes
The human is estimated to have at least 50,000 expressed genes (gene loci). Some information is available concerning about 5000 of these gene loci and about 1900 have been mapped, i.e., assigned to specific chromosomes (and in most instances particular chromosome regions). Progress has been achieved by a combination of physical mapping (e.g., study of somatic cell hybrids and chromosomal in situ hybridization) and genetic mapping (e.g., genetic linkage studies). New methods for both physical and genetic mapping are expanding the armamentarium. The usefulness of the mapping information is already evident; the spin-off from the Human Genome Project (HGP) begins immediately. The complete nucleotide sequence is the ultimate map of the human genome. Sequencing, although already under way for limited segments of the genome, will await further progress in gene mapping, and in particular creation of contig maps for each chromosome. Meanwhile the technology of sequencing and sequence information handling will be developed. It is argued that the HGP is a new form of coordinated, interdisciplinary science; that its primary objective must be seen as the creation of a tool for biomedical research--a source book that will be the basis of study of variation and function for a long time; that the impact on scientist training will be salutary by relieving graduate students of useless drudgery and by training scientists competent in both molecular genetics and computational science; and that the funding of the HGP will have an insignificant negative effect on science funding generally, and indeed may have a beneficial effect through economy of scale and a focusing of attention on the excitement of biology and medical science.
Read moreBiobanking: sample acquisition and quality assurance for ‘omics’ research
Biobanking: sample acquisition and quality assurance for ‘omics’ research