Learning from Structured Data with High Dimensional Structured Input and Output Domain

Fei, Hongliang

dc.contributor.advisor	Huan, Jun
dc.contributor.author	Fei, Hongliang
dc.date.accessioned	2012-11-26T22:35:58Z
dc.date.available	2012-11-26T22:35:58Z
dc.date.issued	2012-08-31
dc.date.submitted	2012
dc.identifier.other	http://dissertations.umi.com/ku:12438
dc.identifier.uri	http://hdl.handle.net/1808/10466
dc.description.abstract	Structured data is accumulated rapidly in many applications, e.g. Bioinformatics, Cheminformatics, social network analysis, natural language processing and text mining. Designing and analyzing algorithms for handling these large collections of structured data has received significant interests in data mining and machine learning communities, both in the input and output domain. However, it is nontrivial to adopt traditional machine learning algorithms, e.g. SVM, linear regression to structured data. For one thing, the structural information in the input domain and output domain is ignored if applying the normal algorithms to structured data. For another, the major challenge in learning from many high-dimensional structured data is that input/output domain can contain tens of thousands even larger number of features and labels. With the high dimensional structured input space and/or structured output space, learning a low dimensional and consistent structured predictive function is important for both robustness and interpretability of the model. In this dissertation, we will present a few machine learning models that learn from the data with structured input features and structured output tasks. For learning from the data with structured input features, I have developed structured sparse boosting for graph classification, structured joint sparse PCA for anomaly detection and localization. Besides learning from structured input, I also investigated the interplay between structured input and output under the context of multi-task learning. In particular, I designed a multi-task learning algorithms that performs structured feature selection & task relationship Inference. We will demonstrate the applications of these structured models on subgraph based graph classification, networked data stream anomaly detection/localization, multiple cancer type prediction, neuron activity prediction and social behavior prediction. Finally, through my intern work at IBM T.J. Watson Research, I will demonstrate how to leverage structural information from mobile data (e.g. call detail record and GPS data) to derive important places from people's daily life for transit optimization and urban planning.
dc.format.extent	185 pages
dc.language.iso	en
dc.publisher	University of Kansas
dc.rights	This item is protected by copyright and unless otherwise specified the copyright of this thesis/dissertation is held by the author.
dc.subject	Computer science
dc.subject	Information science
dc.subject	Anomaly detection
dc.subject	Classification
dc.subject	Data mining
dc.subject	Machine learning
dc.subject	Structrual sparsity
dc.subject	Structured data
dc.title	Learning from Structured Data with High Dimensional Structured Input and Output Domain
dc.type	Dissertation
dc.contributor.cmtemember	Luo, Bo
dc.contributor.cmtemember	Potetz, Brian
dc.contributor.cmtemember	Agah, Arvin
dc.contributor.cmtemember	Xu, Hongguo
dc.thesis.degreeDiscipline	Electrical Engineering & Computer Science
dc.thesis.degreeLevel	Ph.D.
kusw.oastatus	na
kusw.oapolicy	This item does not meet KU Open Access policy criteria.
kusw.bibid	8085805
dc.rights.accessrights	openAccess

Files in this item

Name:: Fei_ku_0099D_12438_DATA_1.pdf
Size:: 5.929Mb
Format:: PDF

View/Open

This item appears in the following Collection(s)

The University of Kansas prohibits discrimination on the basis of race, color, ethnicity, religion, sex, national origin, age, ancestry, disability, status as a veteran, sexual orientation, marital status, parental status, gender identity, gender expression and genetic information in the University’s programs and activities. The following person has been designated to handle inquiries regarding the non-discrimination policies: Director of the Office of Institutional Opportunity and Access, IOA@ku.edu, 1246 W. Campus Road, Room 153A, Lawrence, KS, 66045, (785)864-6414, 711 TTY.