- Conference Article
- 10.1109/bigdata66926.2025.11400897
SegmentAndClassify: A Two-Stage Model for Large Scale Named Entity Recognition
- Dec 08, 2025
- Bin Li
Publications from 2021 to 2026
Showing 10 of 26 papers
SegmentAndClassify: A Two-Stage Model for Large Scale Named Entity Recognition
Towards Equitable Community-Industry Collaborations: Understanding the Experiences of Nonprofits' Collaborations with Tech Companies
Community-based partnerships are essential to creating inclusive and equitable technologies and design practices. Though recent scholarship in HCI focuses on equitable design practices, there is less focus on understanding the experiences of community-based nonprofit organizations (CBOs) when partnering with technology companies. In this paper, we focus on understanding the perspectives of CBOs by answering the following research question: What are the experiences of CBOs that have collaborated with technology companies? Through a series of design workshops with 18 participants who work at community-based nonprofits that have collaborated with technology firms, we identified four elements of community-industry collaborations that collectively shape the overall experience: divergences in cultural and organizational norms, ''setting the table,'' project relationship dynamics, and affective qualities. We conclude by discussing the power structures that impact community-industry collaboration and suggest reflective practices to guide equitable collaborations between CBOs and tech companies.
Read moreBreaking XOR Arbiter PUFs With Chosen Challenge Attack
The XOR Arbiter PUF was introduced as a strong PUF in 2007 and was broken in 2015 by a Machine Learning (ML) attack, which allows the underlying Arbiter PUFs to be modeled individually by exploiting reliability information of the measured responses. To mitigate the reliability-based attacks, state-of-the-art understanding shows that the reliability of individual Arbiter PUFs and the overall XOR Arbiter PUF can be boosted to an arbitrarily high level, thus rendering all known reliability-based ML attacks infeasible; alternatively, an access control interface around the XOR Arbiter PUF can prevent the same challenge-response pairs from being accessed repeatedly, thus eliminating the leakage of reliability information. We show that, for the first time, a perfectly reliable XOR Arbiter PUF can be successfully attacked in a divide-and-conquer manner, meaning each underlying Arbiter PUF in an XOR Arbiter PUF can be attacked individually. This allows us to attack large XOR Arbiter PUFs efficiently, even without reliability information or any side-channel information. Our key insight is that, instead of reliability information, the responses of highly correlated challenges also reveal how close the responses are to the response decision boundary. This leads to a <italic xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">chosen challenge attack</i> on XOR Arbiter PUFs by carefully choosing correlated challenges to measure and aggregate the collected information. We validate our attack by using PUF simulation, as well as an XOR Arbiter PUF implemented on FPGA. We also demonstrate that our chosen challenge methodology is compatible with the state-of-the-art combined gradient-based multi-objective optimization attack. Finally, we discuss an effective countermeasure that can prevent our attack but with a relatively large area overhead compared to the PUF itself.
Read moreComposable Renderer Services: A Study of Architectures for Dynamic View Transclusion in Distributed Web Applications
In large scale enterprise environments, user experi-ence (UX) is delivered through a suite of distributed web applications, each responsible for the various aspects of the overall experience. While this modular approach enhances maintainabil-ity and scalability, it poses challenges while integrating common cohesive UX, present horizontally, across various views of the system. For instance, certain UX elements and functions, like promotional coupons, need to be consistently integrated across disparate views that are powered by varying technology stacks, necessitating a mechanism for seamless transclusion and delivery of such UX content. This paper addresses the need for such a system that enables integration of such UX, horizontally, by introducing a novel composable renderer service. This approach allows for the dynamic transclusion of views rendered by such remote applications into a host application, effectively rendering cohesive UX. The paper compares multiple architectures that house such renderer services. They can be composed with any other application, while preserving the benefits of distributed application design.
Read moreFrom Word Embeddings to Knowledge Graph Embeddings
Word embedding techniques have been developed to assign words to vectors in a vector space. One of the earliest such methods was word2vec, published in 2013 – and embeddings have gathered a tremendous uptake in the natural language processing community since then. Since RDF2vec is based on word2vec, we take a closer look at word2vec in this chapter. We explain how word2vec has been developed to represent words as vectors, and we discuss how this approach can be adapted to knowledge graphs by performing random graph walks, yielding the basic version of RDF2vec. We explain the CBOW and SkipGram variants of basic RDF2vec, revisiting the node classification tasks used in Chap. 1 .
Read moreFinite-sum smooth optimization with SARAH
We introduce NC-SARAH for non-convex optimization as a practical modified version of the original SARAH algorithm that was developed for convex optimization. NC-SARAH is the first to achieve two crucial performance properties at the same time—allowing flexible minibatch sizes and large step sizes to achieve fast convergence in practice as verified by experiments. NC-SARAH has a close to optimal asymptotic convergence rate equal to existing prior variants of SARAH called SPIDER and SpiderBoost that either use an order of magnitude smaller step size or a fixed minibatch size. For convex optimization, we propose SARAH++ with sublinear convergence for general convex and linear convergence for strongly convex problems; and we provide a practical version for which numerical experiments on various datasets show an improved performance.
Read moreMachine Learning in Finance
The finance industry is constantly faced with an ever evolving set of challenges including credit card fraud, identity theft, network intrusion, money laundering, human trafficking, and illegal sales of firearms. There are also newly emerging threats such as fake news in financial media that can lead to distortions in trading strategies and investment decisions. In addition, traditional problems such as customer analytics, forecasting, and recommendations take on a unique flavor when applied to financial data. A number of new ideas are emerging to tackle all these problems including semi-supervised learning methods, deep learning algorithms, network/graph based solutions as well as linguistic approaches. These methods must often be able to work in real-time and be able handle large volumes of data. The purpose of this workshop is to bring together researchers and practitioners to discuss both the problems faced by the financial industry and potential solutions. We have invited regular papers, positional papers and extended abstracts of work in progress. We have also encouraged short papers from financial industry practitioners that introduce domain specific problems and challenges to academic researchers. This event is the fourth in a sequence of finance related workshops we have organized at KDD since 2017.
Read moreAdaReNet: Adaptive Reweighted Semi-supervised Active Learning to Accelerate Label Acquisition
Data scarcity and quality pose significant challenges to supervised learning. The process of generating informative annotations can be time-consuming and often requires high domain expertise. Active and semi-supervised learning methods can reduce labeling effort by either automatically expanding the training set or by selecting the most informative examples to request domain expert annotation. As most selection methods are heuristic, the performance varies widely across datasets and tasks. Bootstrapping approaches such as self-training can result in negative effects due to the addition of incorrectly pseudo-labeled instances. In this work, we take a holistic approach to label acquisition and consider the expansion of clean and pseudo-labeled subsets jointly. To address the challenge of producing high-quality pseudo-labels, we introduce a collaborative teacher-student framework, where the teacher, termed AdaReNet, learns a data-driven curriculum. Experimental results on several natural language processing (NLP) tasks demonstrate that the proposed framework outperforms baselines.
Read morePrice Salience and Product Choice
We analyze a large-scale field experiment on StubHub.com and show that disclosing fees upfront reduces both the quantity and quality of purchases.
Read moreEarly prediction of quality of service using interface-level metrics, code-level metrics, and antipatterns