- Supplementary Content
- 10.2139/ssrn.6036657
Legal Alignment for Safe and Ethical AI
- Jan 01, 2026
- SSRN Electronic Journal
- Noam Kolt + 16 more +16
Publications from 2021 to 2026
Showing 10 of 34 papers
Legal Alignment for Safe and Ethical AI
Flow Autoencoders are Effective Protein Tokenizers
Abstract Protein structure tokenizers enable the creation of multimodal models of protein structure, sequence, and function. Current approaches to protein structure tok-enization rely on bespoke components that are invariant to spatial symmetries, but that are challenging to optimize and scale. We present Kanzi, a flow-based tokenizer for tokenization and generation of protein structures. Kanzi consists of a diffusion autoencoder trained with a flow matching loss. We show that this approach simplifies several aspects of protein structure tokenizers: frame-based representations can be replaced with global coordinates, complex losses are replaced with a single flow matching loss, and SE(3)-invariant attention operations can be replaced with standard attention. We find that these changes stabilize the training of parameter-efficient models that outperform existing to- kenizers on reconstruction metrics at a fraction of the model size and training cost. An autoregressive model trained with Kanzi outperforms similar generative models that operate over tokens, although it does not yet match the performance of state-of-the-art continuous diffusion models. Code is available here: https://github.com/rdilip/kanzi/.
Read moreDV365: Extremely Long User History Modeling at Instagram
Long user history is highly valuable signal for recommendation systems, but effectively incorporating it often comes with high cost in terms of data center power consumption and GPU. In this work, we chose offline embedding over end-to-end sequence length optimization methods to enable extremely long user sequence modeling as a cost-effective solution, and propose a new user embedding learning strategy, multi-slicing and summarization, that generates highly generalizable user representation of user's long-term stable interest. History length we encoded in this embedding is up to 70,000 and on average 40,000. This embedding, named as DV365, is proven highly incremental on top of advanced attentive user sequence models deployed in Instagram. Produced by a single upstream foundational model, it is launched in 15 different models across Instagram and Threads with significant impact, and has been production battle-proven for >1 year since our first launch.
Read moreDeep Learning-Driven Dynamic Clustering for Intelligent Customer Segmentation
E-commerce businesses acquire considerable competitive edge with successful segmentation of customers, allowing for focused marketing campaigns for particular customer groups. This study suggests a new deep learning-based dynamic clustering method that addresses the shortcomings of traditional clustering techniques in dealing with high-dimensional transaction data. Our approach combines autoencoders for non-linear feature extraction and Gaussian Mixture Models (GMM) for probabilistic clustering in feature space with application to the Online Retail Dataset. Experimental results show significant clustering quality improvement in terms of Silhouette Score (0.392 for GMM in latent space) and Davies-Bouldin Index (0.914 for GMM), while enabling business decision-making with actionable customer insights. Performance measurements show that autoencoders capture non-linear relationships more effectively than PCA, even though they use more computational resources (0.367s vs. 0.025s). Our comparison with K-Means, DBSCAN, and Hierarchical Clustering identifies competitive advantages and trade-offs, particularly in dealing with intricate, high-dimensional data. The suggested framework adds to customer analytics by providing an efficient, scalable clustering framework that improves segmentation quality and facilitates business growth strategies. This research also emphasizes the capability of deep learning in solving real-world data segmentation problems, laying the groundwork for future optimization studies.
Read moreSpatial reasoning via recurrent neural dynamics in mouse retrosplenial cortex
From visual perception to language, sensory stimuli change their meaning depending on previous experience. Recurrent neural dynamics can interpret stimuli based on externally cued context, but it is unknown whether they can compute and employ internal hypotheses to resolve ambiguities. Here we show that mouse retrosplenial cortex (RSC) can form several hypotheses over time and perform spatial reasoning through recurrent dynamics. In our task, mice navigated using ambiguous landmarks that are identified through their mutual spatial relationship, requiring sequential refinement of hypotheses. Neurons in RSC and in artificial neural networks encoded mixtures of hypotheses, location and sensory information, and were constrained by robust low-dimensional dynamics. RSC encoded hypotheses as locations in activity space with divergent trajectories for identical sensory inputs, enabling their correct interpretation. Our results indicate that interactions between internal hypotheses and external sensory data in recurrent circuits can provide a substrate for complex sequential cognitive reasoning.
Read moreRL-Finetuning of OpenAI o1-mini to Enhance Biomedical Reasoning
Abstract Recent breakthroughs in advanced reasoning large language models (LLMs), such as OpenAI’s o1, have achieved impressive results in domains like math and coding. However, it’s not clear how much this type of reasoning helps in solving biomedical problems that involve more domain specialized knowledge and open-ended reasoning. Across two biomedical domains—gene characterization and small molecule property prediction—we find that the commercially available o1-mini model does not consistently outperform non-reasoning LLMs like GPT-4o. This motivated us to explore how much we can improve o1-mini’s biomedical reasoning through reinforcement learning (RL) finetuning. We show that RL finetuning of o1-mini results in large improvements in performance on gene classification, where it surprisingly outperformed domain-specific state-of-the-art models on some tasks. The results are mixed for small molecule prediction, suggesting that chemical reasoning could be more challenging for LLMs. We conclude with a discussion of the challenges and takeaways from this initial exploration of RL finetuning reasoning models for biomedical tasks.
Read moreArtVLM: Attribute Recognition Through Vision-Based Prefix Language Modeling
A Note on the Maximum Number of k-Powers in a Finite Word
A power is a concatenation of $k$ copies of a word $u$, for a positive integer $k$; the power is also called a $k$-power and $k$ is its exponent. We prove that for any $k \ge 2$, the maximum number of different non-empty $k$-power factors in a word of length $n$ is between $\frac{n}{k-1}-\Theta(\sqrt{n})$ and $\frac{n-1}{k-1}$. We also show that the maximum number of different non-empty power factors of exponent at least 2 in a length-$n$ word is at most $n-1$. Both upper bounds generalize the recent upper bound of $n-1$ on the maximum number of different square factors in a length-$n$ word by Brlek and Li (2022).
Read moreGeneralized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images
Capturing the diversity of people in images is challenging: recent literature tends to focus on diversifying one or two attributes, requiring expensive attribute labels or building classifiers. We introduce a diverse people image ranking method which more flexibly aligns with human notions of people diversity in a less prescriptive, label-free manner. The Perception-Aligned Text-derived Human representation Space (PATHS) aims to capture all or many relevant features of people-related diversity, and, when used as the representation space in the standard Maximal Marginal Relevance (MMR) ranking algorithm [7] , is better able to surface a range of types of people-related diversity (e.g. disability, cultural attire). PATHS is created in two stages. First, a text-guided approach is used to extract a person-diversity representation from a pre-trained imagetext model. Then this representation is fine-tuned on perception judgments from human annotators so that it captures the aspects of people-related similarity that humans find most salient. Empirical results show that the PATHS method achieves diversity better than baseline methods, according to side-by-side ratings from human annotators.
Read moreTopological Embedding of Human Brain Networks with Applications to Dynamics of Temporal Lobe Epilepsy
We introduce a novel, data-driven topological data analysis (TDA) approach for embedding brain networks into a lower-dimensional space in quantifying the dynamics of temporal lobe epilepsy (TLE) obtained from resting-state functional magnetic resonance imaging (rs-fMRI). This embedding facilitates the orthogonal projection of 0D and 1D topological features, allowing for the visualization and modeling of the dynamics of functional human brain networks in a resting state. We then quantify the topological disparities between networks to determine the coordinates for embedding. This framework enables us to conduct a coherent statistical inference within the embedded space. Our results indicate that brain network topology in TLE patients exhibits increased rigidity in 0D topology but more rapid flections compared to that of normal controls in 1D topology.
Read more