Text Mining Innovation for Business
Text Mining Innovation for Business
Ela Pustulka and Thomas Hanne
Abstract This chapter reflects on the business innovation supported by developing text mining solutions to meet the business needs communicated by Swiss compa- nies. Two related projects from different industries and with different challenges are discussed in order to identify common procedures and methodologies that can be used. One of the partners, in the gig work sector, offers a platform solution for employee recruitment for temporary work. The work assessment is performed using short reviews for which a method for sentiment assessment based on machine learning has been developed. The other partner, in the financial advice sector, oper- ates an information extraction service for business documents, including insurance policies. This requires automation in the extraction of structured information from pdf-files. The common path to innovation in such projects includes business process modeling and the implementation of novel technological solutions, including text mining techniques.
Keywords Digitalization · Text mining · BPMN · Innovation · Anonymization · Gig work · Insurance brokerage
1 Introduction
Digital transformation is one of the most important business drivers nowadays [1]. Two concepts come into play in this area: digitalization and digitization. Roughly speaking, digitization is the underlying change in the way we store and use infor- mation, where we replace analog/physical storage and data handling with digital information formats and data flows. Digitalization, on the other hand, deals with
E. Pustulka (B) · T. Hanne Institute for Information Systems, School of Business, FHNW University of Applied Sciences and Arts Northwestern Switzerland, Riggenbachstrasse 16, 4600 Olten, Switzerland e-mail: [email protected]
T. Hanne e-mail: [email protected]
© Springer Nature Switzerland AG 2021 R. Dornberger (ed.), New Trends in Business Information Systems and Technology, Studies in Systems, Decision and Control 294, https://doi.org/10.1007/978-3-030-48332-6_4
49
50 E. Pustulka and T. Hanne
processes that can be completely new and have little to do with previous processes that were the traditional way of doing business on paper or in other physical ways.
Both digitization and digitalization lead to profound changes in the way business and society work. We see both processes as both sources and outcomes of innova- tion. In digital transformation, the process of change usually takes place organically within companies and originates within the minds of business leaders who shape their business in such a way that they take advantage of the opportunities that arise. Companies then often use external technical expertise to guide the development of new business solutions [2, 3]. In this chapter, we focus on how digital transforma- tion benefits from novel solutions based on text mining, which we develop with our business partners.
The path to innovation is shaped by the business partners who state their needs. Based on our experiences, the following steps in such projects are usually useful and should be carried out in close collaboration with the project partners:
1. creating Business Process Modeling and Notation (BPMN) models [4] 2. obtaining datasets from the business 3. preliminary experiments to assess feasibility, based on the available data 4. designing an anonymization procedure (if needed) 5. obtaining more data 6. prototyping and assessment 7. refinement 8. business adaptation and use.
This approach combines the process with the data that reflect the common percep- tion of text mining and data science defined as a “broader discipline that combines knowledge from information technology and knowledge from management sciences to improve and run operational processes” [5], p. 811.
However, contrary to the assessment made by Aalst and Damiani [5] who expect companies to have many processes with logs that can be mined, we work with small or medium-sized enterprises (SMEs) that have often not documented their processes anddonothavesuchlogs,andthuswedocumenttheprocessesmanually.Wecarryout a business process analysis to understand the business requirements and to develop a shared vision of the future with our partners. Visualizing business processes with BPMN is useful because it provides a level of clarity and a communication basis that our partners can relate to. In one case, the analysis led the company to realize that their processes had gaps that were not under their control, which is now being addressed [6].
The second step in our approach is to gain access to datasets for text mining. After initial experiments, we often find that there are not enough data for further analysis, either in terms of volume or data complexity, and we ask for more. Since data privacy needs to be guaranteed [7], we design anonymization solutions that give us clean data that comply with the law. Then, in close collaboration with our partner, we develop prototypes and evaluate them together to see if they meet business needs. Practice shows that a viable solution has to provide high accuracy, with around 90% being
Text Mining Innovation for Business 51
sufficient. Finally, the companies take over the prototypes and integrate them into products that offer business value.
Our contribution is the description of the innovation processes and their outcomes, which lead to two business innovations. One is a text mining solution for the gig work sector, which uses text mining for sentiment analysis. The other is ongoing work on the development of an information extraction module for a platform that supports insurance brokerage. In both cases, the development of individual solutions was necessary because there was no suitable standard software available.
In addition to discussing some specific results of the two projects, this chapter points out the partly common procedures and methodologies employed in the two considered cases, which come from different industries and focus on different chal- lenges. In Sect. 2, BPMN is discussed regarding its usefulness in innovation-oriented projects, but also to introduce the reader to the background of the two projects and the respective business domains. In particular, we present the developed BPMN models. Section 3 deals with general aspects of text mining in the considered projects. Three specific application areas of text mining are discussed: the anonymization of documents, sentiment analysis of user comments, and information extraction from unstructured text files. Section 4 discusses our contributions and concludes.
2 BPMN Analysis for Applied Research Projects
We use Business Process Modeling and Notation (BPMN) to help us understand the business scenario and requirements [4]. The graphical notation improves the general understanding of the current processes and errors in our judgement can be clearly spotted and corrected. We work in two business areas: evaluation of gig work jobs, see Sect. 2.1 and information extraction serving the needs of financial service advisors (insurance brokers), see Sect. 2.2.
2.1 Gig Work Platform Processes
Greenwood et al. [8], p. 27 define gig-economy platforms as “digital, service-based, on-demand platforms that enable flexible work arrangements”. Our business partner in the gig work sector has approximately 150 employees in Europe. The company operatesapart-electronicHRrecruitmentplatformandcanfindemployeesfortempo- rary work assignments within a few hours of receiving an order from a potential employer. The company uses traditional human-intensive processes to acquire new business (employers) and mostly electronic processes to recruit workers and manage the employment process. We initiated a joint research project focusing initially on improving electronic worker recruitment and retention. Our research looks at the entire business process. We identified some issues based on the priorities specified
52 E. Pustulka and T. Hanne
Fig. 1 Simplified illustration from [6]. Gig work process overview showing the worker lane, the platform company lane, and the employer lane. Processes are numbered 1–6, with letters beside the number, to distinguish between activities performed by the participants. Process 1, REGISTER, is visible in the left part of the figure in all three lanes (1A to 1C). Process 2, HIRE, follows, as 2A– 2C. Process 3, GIG WORK, is outside the platform. Process 4, GIG COMPLETION/REPORTING, involves all three lanes and is followed by Process 5, ANALYSE RATINGS (5A to 5C). Finally, Process 6, ACCOUNTING, closes the transaction. All participants are modeled in one pool because they share one marketplace
by the project partner (sentiment analysis and topic analysis of job reviews) and some future work (automated analysis of CVs and assignment of job labels to them).
The company gave us a demo of the platform system and granted access to their Wiki, which included platform documentation [6] in the form of a matrix showing the system functions, both implemented and planned. In a short modeling session we translated those functions into a BPMN model, see Fig. 1. This model was discussed with the business partner, corrected, and frozen. In this case, the process is so closely tied to the underlying platform that it is not changed until a new software version is released, which is rare (from several months to a year).
Figure 1 shows an outline of the platform process which connects employers to employees. The process can logically be divided into interactions with employers and employees, leading to three lanes in the diagram. From left to right, we see process steps that start with the initiation of the business relationship (registration, step 1), followed by hiring (step 2), job execution (work, step 3), gig completion/reporting (step 4), job performance analysis (ratings, step 5) and accounting (step 6). Steps 2 to 6 are repeated for each new temporary assignment for which the prospective employer wishes to recruit.
TextminingwillsupporttheProcess5B,Analyzeratings.Afteragig,bothworkers and employers rate each other. They provide a star rating (1–4 stars, with 1 being poor and 4 excellent) and, if the rating is low (1–2 stars), they are required to enter a comment explaining their rating. Comments can also be provided at high rating. Currently, the platform company only communicates the star rating to the worker and to the company (Processes 5A and 5C), together with the confirmed number of hours worked. The job reviews, however, are not used. The business partner wants
Text Mining Innovation for Business 53
to communicate them to the parties involved and is in the process of deciding how best to do this, as this needs to be part of the platform system. The project supports this process by prototyping sentiment and topic analysis tools (both based on text mining techniques) and proposing a sentiment dashboard, as detailed in Sect. 3.3.
2.2 An Information Extraction Service for Insurance Brokers
The second innovation project discussed in this chapter involves the provision of financial advice services such as decisions regarding insurance contracts. Currently, insurance brokers receive scanned documents from customers and take several minutes per document to manually extract the most relevant information from the document into the software they use to prepare a recommendation regarding insurance-relevant requirements. The scenario is an extension of ideas already in operation (Evia, “expert virtual insurance agent” [9]) and of the trend to automate business processes and replace them with Robotic Process Automation [10]. The business scenario includes three parties: the customer, the broker, and the company providing the extraction service. The envisaged BPMN process is outlined in Fig. 2 (customer interaction not shown). The Extraction Service extracts information from pdfs and produces structured data. This service is being implemented using text mining.
Let us consider insurance policies as documents and focus on those issued in Switzerland. There are up to ten large insurance companies and around 15 types of insurance (vehicle, health, legal, etc.). Each of the insurance types encompasses several possible types of coverage and related data (sum covered, period covered,
Fig. 2 An overview of the planned interactions between financial advice brokers and the Extraction Service (ES). The insurance broker receives the paperwork from the customer, sends the documents to the ES and obtains structured information which can be used to prepare a recommendation for the customer. After the billing, the process ends
54 E. Pustulka and T. Hanne
exclusions, etc.). A policy covers only a selection of clauses and each of these clauses typically has partial costs related to the coverage and the insured sum that are relevant. It is possible to use information extraction technologies based on pdf annotation, such as smartFIX [11], and create an extraction schema for each policy individually. However, the analysis and full implementation take around eight days per policy. Additionalcomplicationsintheuseofcurrentinformationextractiontoolscomefrom the insertion of advertising material and format changes that make the tools unstable. The goal of the project is to design a tool for automated information extraction from pdf documents. The tool will support the business process shown in Fig. 2. Text mining will support information extraction as detailed in Sect. 3.4.
3 Text Mining Challenges and Solutions in Applied Research Projects
3.1 Tasks and Challenges in Text Mining
Weiss et al. [12] draw the distinction between data mining and text mining as “num- bers vs. text”. In text mining, the text is transformed into numbers and then data mining procedures can be applied to a numerical representation. Typical applica- tions of text mining are classification and prediction. In a classification setting, the text is to be divided into “equivalence classes” with shared characteristics. In the prediction setting, statements are to be made about the future, based on the past. Text mining is used in information retrieval [13]. It is also used to cluster and organize document collections, and support browsing [14]. Another application is informa- tion extraction, i.e. converting text documents to structured documents that can be searched and compared using database methods, or in data integration, e.g. [15]. Finally, in a predictive scenario, text mining can predict user needs during search or browsing and support the user with suggestions.
As the applications described in this chapter are located in the area of human resource management and financial advice, this section only covers the recent use of text mining in these two domains. Piazza and Strohmeier [16] review the uses of data mining in human resource management. In their opinion, the main success factors in this area are related to the functional dimension (staffing, development, performance management, and compensation), method dimension (algorithms), data provision, information system provision, user support, and ethical and legal awareness. Their review of 100 research contributions finds that 90% of work used data mining and 10% used text mining or web mining. Thus, only 10% of papers reviewed in [16] make provisions for data privacy, which may be implemented via aggregation or de- personalization (anonymization). Our scenarios require anonymization, as detailed in Sect. 3.2.
Text mining is used extensively in the financial sector. Kumar and Ravi [14] describemethodsandapplications.TheyfindthatSVM(supportvectormachines)are
Text Mining Innovation for Business 55
the most popular method, followed by naïve Bayes (NB), k-nearest neighbors (k-NN) and decision trees. Financial applications include foreign exchange prediction, stock market prediction, customer relationship management (CRM), and cybersecurity. In the areas most related to our projects, they only mention the work of Ghani et al. [17] on product attribute extraction. Other work in the area includes [18] describing a system called Midas that uses information extraction, entity resolution, mapping and fusion to integrate public data and support systemic risk analysis in the financial market. Midas can help discover suspicious activities such as co-lending, and risk at the company level, including relationships with other companies, key executives and aggregated financial data. However, systems like Midas are too complex for an SME and they need to be tuned to fit the business scenario. Tuning via expert user interaction is also required in newer information extraction systems [19] and may be a barrier to system adoption.
Text mining scenarios are a challenge to business for several reasons:
• Data may need to be anonymized (legal requirement). • Large amounts of data are required to build statistical machine learning models
that guarantee the required quality of results, and many SMEs may not have enough data to use data mining.
• As text mining is constantly evolving, being aware of the latest developments in research (including new methods or open source software) is a prerequisite.
• Several rounds of prototyping are needed to develop solutions of the required quality, as published methods do not necessarily deliver the best results on company data and need to be extended, tuned, or combined with other published methods.
• Domain specific adjustments are required and result in improved quality. In the projects described in this chapter, this involves the language, specific vocabu- lary, and other text characteristics (short utterances or even lists). In the senti- ment analysis described in Sect. 3.3, there are three languages (German, French, and English) with very short textual comments. In information extraction (see Sect. 3.4) the situation is similar, as we have financial documents in German, French, English, and Italian, including insurance policies, tax statements, bank statements, and ownership certificates.
3.2 Data Anonymization
Text mining requires considerable amounts of data to achieve the quality acceptable for business use. Sensitive personal documents, such as insurance policies or bank statements, are a challenge because they cannot be used directly in text mining due to data protection regulations [7]. In our work (in both projects), we started with a simple approach that was appropriate in the context of sentiment analysis, and developed it further to anonymize financial documents. Providing anonymization software to the business partner was the best way to ensure that we did not handle private data, as the original data remained with its owners.
56 E. Pustulka and T. Hanne
Table 1 Pseudo-code of the preprocessing and anonymization algorithm, used with each file, called pdffile. Tokenize splits on space. Intersection is the set intersection
In the project with the gig work company, the amount of text that had to be anonymized consisted of only a few thousand short text fragments. The data had random IDs and we only had to remove names, dates and, phone numbers from the text and comments written by the employers and the employees. We used regular expressions (REs) and a dictionary of the thousand most common Swiss names from a public website in order to identify the names in the documents. We also searched for expressions starting with “Mr”, “Mrs”, “Customer” and their equivalents in three languages. However, this approach did not guarantee high quality, which made us develop a more complex solution designed for financial data.
The anonymization algorithm we developed subsequently relies on the existing information sources available in Switzerland. First, we use an address register provided by the Swiss Post [20], out of which we extract valid street names, here called street. Second, we use the register of the most common Swiss given names [21], here called firstnames. The algorithm is shown in Table 1. It first loads the streets and first names into sets (lines 1 and 2). It then traverses the document and uses Pdfminer [22] to extract the bounding boxes (line 5). The text extracted from these is turned into a set of words and set intersection delivers the tokens that are to be anonymized (lines 8 and 9). These tokens are then processed by REs in accordance with customer requirements (lines 11–16).
3.3 Sentiment Analysis
The analysis of the gig work platform processes and the discussions with the project partners led to the identification of the need for a sentiment analysis of job reviews written by the workers and their employers. As mentioned above, reviews carry a star rating with 1–2 stars seen as “negative” and 3–4 stars as “positive”. However,
Text Mining Innovation for Business 57
the star ratings are often not consistent with the textual content of the comments [23], as pointed out by the company. This led to an effort to automatically assign a sentiment value, with classes defined differently: “negative” are the statements containing criticism and “other” are the positive or neutral ones, no matter what the star rating was. “Negative” statements are of interest as they point to problems and can help improve the business and ensure higher customer retention.
A summary of the results [23] is shown in Table 2. In prototyping, manually assigned sentiment labels on 963 job reviews were used, 26% of which were labeled “negative”. Machine learning was based on Scikit-learn [24]. The highest accuracy of automated sentiment assignment was 86%, with logistic regression and SVM performing similarly (shown in bold). As the data set is unbalanced, we also show the Matthews correlation coefficient [25].
This initial analysis was discussed with our business partner and led to further development, with additional data annotation and improved quality as shown in Table 3. The resulting accuracy of over 90% is acceptable in the business scenario. Figure 3 shows the learning curve that flattens out as more data is being added to the model.
The next step in the project was to position the use of sentiment analysis prediction in the business. The gig work company operates a help desk and a customer care departmentthatmonitordailyactivityandlookattrends.Similarly,salesmanagement regularly reviews the quality. In the project, a sentiment display was prototyped, part of which is shown in Fig. 4.
Table 2 Sentiment analysis results as reported in [23] with 963 rows of data. SVM is a support vectorclassifier,k-NNstandsfork-nearestneighbors,Treeisadecisiontree,LRislogisticregression and RF is random forest. Accuracy, Matthews’ coefficient and the F1 measure are shown for these methods
SVM C = 1 k-NN Tree LR RF Accuracy 0.866 0.836 0.820 0.869 0.839
Matthews 0.634 0.555 0.515 0.646 0.551
F1 0.913 0.893 0.881 0.914 0.897
Table 3 Improved sentiment analysis after further manual sentiment scoring, with 2428 rows of data. Logistic regression performs better than the support vector machine (both in bold). The confusion matrices are shown with the target class containing negative comments in the top left corner
SVM C = 1 k-NN Tree LR RF Accuracy 0.906 0.870 0.862 0.920 0.867
Matthews 0.732 0.622 0.617 0.777 0.613
F1 0.939 0.917 0.910 0.948 0.915
Confusion matrix
[[428, 154], [74, 1772]]
[[358, 224], [91, 1755]]
[[402, 180], [154, 1692]]
[[466, 116], [77, 1769]]
[[353, 229], [93, 1753]]
58 E. Pustulka and T. Hanne
Fig. 3 The learning curve for logistic regression with the larger dataset shows that adding more data does not improve the results dramatically, but a marginal improvement should still be possible
Fig. 4 A prototype of sentiment visualization showing company sentiment towards workers over time. It is possible to select only the ratings with textual comments (tick box top right), select a time range by using the slider, or view the text details and reclassify the comments (i.e. change negative to other or the other way around, not shown)
Ongoing work is addressing the topics appearing in the sentiment analysis. We are testing the use of biterms [26] to identify topics in multilingual text. The idea is to match the topics that are found automatically with worker and company concerns expressed via other channels such as face to face or helpdesk contact.
Text Mining Innovation for Business 59
3.4 Extraction of Structured Information from Unstructured Text
Text mining in the financial advice sector (insurance brokerage) supports the devel- opment of an Extraction Service (ES) shown in Fig. 2. The service will parse pdf documents from customers and deliver structured data for financial advice brokers. Information extraction is an established research area [15] with recent work including [19]. Our approach builds in particular on the work on document similarity, such as recent CV to job matching done by Xu and Barbosa [27], and builds a domain specific solution consisting of the following elements:
• A knowledge base reflecting the internal company schemas and an external dictionary of the terminology used in the area.
• Document anonymization as shown in Table 1. • Document sectioning, based on font size and layout. • Calculating text similarity measures derived from documents and the knowledge
base and using them to improve the sectioning and identify features and algorithms that work well in our context.
• Selecting the features to be used in segmentation and matching and using an optimization algorithm to find the best features and algorithms, inspired by results reported by Fua and Hanson [28].
• Refining our solutions to suit the business need and extending them to further document types.
Based on a large set of anonymized documents from our business partner, we are entering the document analysis phase that will identify potential solutions. In addition, a knowledge base capturing company schemas is to be built.
4 Conclusions
Our work applies the latest research findings in text mining to business scenarios where innovation is needed. Applied research projects shape a new business reality and help companies capture new markets or tame new technologies. These new developments are part of digital transformation as they change business processes and offload manual data processing to artificial intelligence solutions.
We have performed business process analysis and specified the business require- ments. Text mining has already delivered reliable methods for sentiment analysis and document anonymization. We are now working on topic analysis and informa- tion extraction and trying out various methods to find the best combinations of text features and data mining algorithms. Our future work will be to tailor the known solutions to the domain of interest and to use machine learning.
For the considered fields of application (sentiment analysis, document anonymiza- tion, extraction of structured information) and text mining problems in general, there
60 E. Pustulka and T. Hanne
are many methods available but no standard approach will solve all the problems equally well. Usually, it depends on the specific data which approaches work well enough and which do not. Therefore, a comprehensive evaluation of the methods, the individual preprocessing of data and the fine-tuning of the methods is suggested.
Unfortunately, the limited budget of research projects does not allow for extensive explorations of alternative approaches. In addition, time constraints often suggest developing a solution that is sufficiently efficient, but not necessarily optimal.
Acknowledgements Funding was provided by the Swiss Commission for Technology and Innovation CTI (now Innosuisse): 25826.1 PFES-ES and 34604.1 IP-ICT, and by the FHNW.
References
1. Davidovski, V.: Exponential innovation through digital transformation. In: Proceedings of the 3rd International Conference on Applications in Information Technology, pp. 3–5. ACM, New York, NY, USA (2018). https://doi.org/10.1145/3274856.3274858
2. Ivascu, L., Cirjaliu, B., Draghici, A.: Business model for the university-industry collaboration in open innovation. Procedia Econ. Finan. 39, 674–678 (2016). https://doi.org/10.1016/S2212- 5671(16)30288-X
3. Nambisan, S., Lyytinen, K., Majchrzak, A., Song, M.: Digital innovation management: rein- venting innovation management research in a digital world. MIS Q. 41, (2017). https://doi.org/ 10.25300/misq/2017/41.1.01
4. Weske, M.: Business Process Management: Concepts, Languages, Architectures. Springer (2012). https://doi.org/10.1007/978-3-642-28616-2_7
5. Van der Aalst, W., Damiani, E.: Processes meet big data: connecting data science with process science. IEEE Trans. Serv. Comput. 8, 810–819 (2015). https://doi.org/10.1109/tsc.2015.249 3732
6. Pustulka-Hunt, E., Telesko, R., Hanne, T.: Gig work business process improvement. In: 2018 6th International Symposium on Computational and Business Intelligence (ISCBI), pp. 10–15 (2018). https://doi.org/10.1109/iscbi.2018.00013
7. Calder, A.: EU GDPR: A Pocket Guide. IT Governance Publishing, Ely, Cambridgeshire, United Kingdom (2018). https://www.jstor.org/stable/j.ctt1cd0mkw
8. Greenwood, B., Burtch, G., Carnahan, S.: Unknowns of the gig-economy. Commun. ACM 60(7), 27–29 (2017). https://doi.org/10.1145/3097349
9. CNET. Your Next Insurance Agent Will Be a Robot. https://cacm.acm.org/careers/197572- your-next-insurance-agent-will-be-a-robot/fulltext. Accessed April 23, 2019
10. Reich, M., Braasch, T.: Die Revolution der Prozessautomatisierung bei Versicherungsun- ternehmen: Robotic Process Automation (RPA). In: Handbuch Versicherungsmarketing, pp. 291–305. Springer (2019). https://doi.org/10.1007/978-3-662-57755-4_17
11. smartFix. https://www.insiders-technologies.de/home/products/input-management/general- incoming-mail/smart-fix.html. Accessed 23 April 2019
12. Weiss, S.M., Indurkhya, N., Zhang, T.: Fundamentals of Predictive Text Mining. Springer Publishing Company, Incorporated (2016). https://doi.org/10.1007/978-1-84996-226-1
13. Baeza-Yates, R.A., Ribeiro-Neto, B.: Modern Information Retrieval. Addison-Wesley Longman Publishing Co., Inc, Boston, MA, USA (1999)
14. Kumar, B.S., Ravi, V.: A survey of the applications of text mining in financial domain. Know.- Based Syst. 114, 128–147 (2016). https://doi.org/10.1016/j.knosys.2016.10.003
Text Mining Innovation for Business 61
15. Baumgartner, R., Flesca, S., Gottlob, G.: Visual web information extraction with lixto. In: Proceedings of the 27th International Conference on Very Large Data Bases, pp. 119–128. Morgan Kaufmann Publishers Inc. (2001)
16. Piazza, F., Strohmeier, S.: Domain-driven data mining in human resource management: a review. In: Proceedings of the 2011 IEEE 11th International Conference on Data Mining Workshops, pp. 458–465. IEEE Computer Society (2011). https://doi.org/10.1109/icdmw.201 1.68
17. Ghani, R., Probst, K., Liu, Y., Krema, M., Fano, A.: Text mining for product attribute extraction. SIGKDD Explor. Newsl. 8, 41–48 (2006). https://doi.org/10.1145/1147234.1147241
18. Burdick, D., Hernández, M.A., Ho, C.T.H., Koutrika, G., Krishnamurthy, R., Popa, L., Stanoi, I., Vaithyanathan, S., Das, S.R.: Extracting, linking and integrating data from public sources: a financial case study. IEEE Data Eng. Bull. 34, 60–67 (2011). https://doi.org/10.2139/ssrn.266 6384
19. Staar, P.W.J., Dolfi, M., Auer, C., Bekas, C.: Corpus conversion service: a machine learning platform to ingest documents at scale. In: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pp. 774–782. ACM (2018). https:// doi.org/10.1145/3219819.3219834
20. Die Post. https://www.post.ch/de/geschaeftlich/themen-a-z/adressen-pflegen-und-geodaten- nutzen/adress-und-geodaten. Accessed 11 April 2019
21. Vornamen in der Schweiz. Vornamen der Bevölkerung nach Geschlecht, Schweiz, 2017. BFS- Nummer: su-t-01.04.00.12. Accessed 23 April 2019
22. Pdfminer. Yusuke Shinyama. Pdfminer is a tool for extracting information from PDF documents. https://github.com/euske/pdfminer. Accessed 23 April 2019
23. Pustulka-Hunt, E., Hanne, T., Blumer, E., Frieder, M.: Multilingual sentiment analysis for a swiss gig. In: 2018 6th International Symposium on Computational and Business Intelligence (ISCBI), pp. 94–98 (2018). https://doi.org/10.1109/iscbi.2018.00028
24. Avila, J.: Scikit-learn cookbook: over 80 recipes for machine learning in python with scikit-learn. Packt Publishing, Birmingham, UK (2017). https://doi.org/10.1214/009053604 000000067
25. Matthews, B.W.: Comparison of the predicted and observed secondary structure of T4 phage lysozyme. Biochim. et Biophys. Acta (BBA)-Protein Structure 405, 442–451 (1975)
26. Yan, X., Guo, J., Lan, Y., Cheng, X.: A biterm topic model for short texts. In: Proceedings of the 22nd International Conference on World Wide Web, pp. 1445–1456. ACM (2013). https:// doi.org/10.1145/2488388.2488514
27. Xu, P., Barbosa, D.: Matching résumés to job descriptions with stacked models. In: Advances in Artificial Intelligence: 31st Canadian Conference on Artificial Intelligence, Canadian AI 2018, Toronto, ON, Canada, May 8–11, 2018, Proceedings 31, pp. 304–309. Springer (2018). https://doi.org/10.1007/978-3-319-89656-4_31
28. Fua, P., Hanson, A.: An optimization framework for feature extraction. Mach. Vis. Appl. 4, 59–87 (1991)
- Text Mining Innovation for Business
- 1 Introduction
- Text Mining Innovation for Business
- 2 BPMN Analysis for Applied Research Projects
- 2.1 Gig Work Platform Processes
- Text Mining Innovation for Business
- 2 BPMN Analysis for Applied Research Projects
- 2.2 An Information Extraction Service for Insurance Brokers
- Text Mining Innovation for Business
- 3 Text Mining Challenges and Solutions in Applied Research Projects
- 3.1 Tasks and Challenges in Text Mining
- Text Mining Innovation for Business
- 3 Text Mining Challenges and Solutions in Applied Research Projects
- 3.2 Data Anonymization
- Text Mining Innovation for Business
- 3 Text Mining Challenges and Solutions in Applied Research Projects
- 3.3 Sentiment Analysis
- Text Mining Innovation for Business
- 3 Text Mining Challenges and Solutions in Applied Research Projects
- 3.4 Extraction of Structured Information from Unstructured Text
- 4 Conclusions
- Text Mining Innovation for Business
- References