Information Gathering and Classification for Collaborative Logistics Decision Making
Collaboration is perceived as a powerful tool for companies to deal with an increasingly demanding global economic environment. Collaboration impact upon the logistics processes is analyzed in this chapter. Logistics naturally implies the concurrency of a set of companies being part of products and services value chain. However, collaborative logistics imposes new challenges that effectively faced by companies and government institutions will produce the expected return. Such challenges relate to effective decision making process involving logistics activities planning, scheduling, control, and coordination at different companies being part of the logistics network. These decision making processes require the effective capability to capture, process, and analyze information on processes, partners, environment, regulations, etc. That information is, of course, dynamic, and it is available in a scattered way all through a variety of sources, periodicity, and formats. Effective and efficient tools for capturing, classifying, processing, and report useful information for supporting the collaborative logistics decision making processes are required. Large amounts of digital text available on the web contain useful information for enabling collaborative logistics. The amount of digital text it is expected to increase significantly in the near future, making the development of data analysis applications an urgent need. Information gathering has been traditionally faced integrating systems and databases at the various institutions and companies participating in the logistics network. This approach has at least two difficulties that have not been solved properly: 1) It requires the data be available in structured databases and stored in terms of the same attributes, and 2) Organizations need to provide access to their systems for obtaining and processing the data. An alternative approach is using data available in the web as input to automatic text classifiers implemented with machine learning techniques. Automatic text classification (or categorization) is defined as the assignment of a Boolean value to each pair (dj,ci) belonging to the set D×C, where D is the domain of documents and C = {c1,...c|C|} is the set of predefined labels. Binary classification is the most simple and widely studied case, in which a document is classified into one of two mutually exclusive categories or classes. Document representation has a high impact on the task of classification. Some elements used for representing documents are: N-grams, single-word, phrases, or logical terms and statements. The vector space model is one of the most widely used models for ad-hoc information retrieval, mainly because of its conceptual simplicity and the appeal of the
Read more