Publications (31)
Dense Captioning with Joint Inference and Visual Context
Linjie Yang, Kevin Tang, Jianchao Yang +1
Dense captioning is a newly emerging computer vision topic for understanding images with dense language descriptions. The goal is to densely detect visual concepts (e.g., objects,…
Prosody leaks into the memories of words
Kevin Tang, Jason A. Shaw
The average predictability (aka informativity) of a word in context has been shown to condition word duration (Seyfarth, 2014). All else being equal, words that tend to occur in mo…
Analysis of LLM as a grammatical feature tagger for African American English
Rahul Porwal, Alice Rozet, Pryce Houck +3
African American English (AAE) presents unique challenges in natural language processing (NLP). This research systematically compares the performance of available NLP models--rule-…
Learning Temporal Embeddings for Complex Video Analysis
Vignesh Ramanathan, Kevin Tang, Greg Mori +1
In this paper, we propose to learn temporal embeddings of video frames for complex video analysis. Large quantities of unlabeled video data can be easily obtained from the Internet…
Modeling Probabilistic Reduction using Information Theory and Naive Discriminative Learning
Anna Stein, Kevin Tang
This study compares probabilistic predictors based on information theory with Naive Discriminative Learning (NDL) predictors in modeling acoustic word duration, focusing on probabi…
Connecting the Persian-speaking World through Transliteration
Rayyan Merchant, Akhilesh Kakolu Ramarao, Kevin Tang
Despite speaking mutually intelligible varieties of the same language, speakers of Tajik Persian, written in a modified Cyrillic alphabet, cannot read Iranian and Afghan texts writ…
Distribution of Si II 6355 Velocities of Type Ia Supernovae and Implications for Asymmetric Explosions
Keto D. Zhang, WeiKang Zheng, Thomas de Jaeger +7
The ejecta velocity is a very important parameter in studying the structure and properties of Type Ia supernovae (SNe Ia). It is also a candidate key parameter in improving the uti…
Investigating the Nature of the Luminous Ambiguous Nuclear Transient ASASSN-17jz
Thomas W. -S. Holoien, Jack M. M. Neustadt, Patrick J. Vallely +31
We present observations of the extremely luminous but ambiguous nuclear transient (ANT) ASASSN-17jz, spanning roughly 1200 days of the object's evolution. ASASSN-17jz was discovere…
The Lick Observatory Supernova Search follow-up program: photometry data release of 70 stripped-envelope supernovae
WeiKang Zheng, Benjamin E. Stahl, Thomas de Jaeger +85
We present BVRI and unfiltered Clear light curves of 70 stripped-envelope supernovae (SESNe), observed between 2003 and 2020, from the Lick Observatory Supernova Search (LOSS) foll…
Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English
Hamid Mojarad, Kevin Tang
Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, and at what level of abstracti…
A Sociolinguistic Analysis of Automatic Speech Recognition Bias in Newcastle English
Dana Serditova, Kevin Tang
Automatic Speech Recognition (ASR) systems are widely used in everyday communication, education, healthcare, and industry, yet their performance remains uneven across speakers, par…
ParsTranslit: Truly Versatile Tajik-Farsi Transliteration
Rayyan Merchant, Kevin Tang
As a digraphic language, the Persian language utilizes two written standards: Perso-Arabic in Afghanistan and Iran, and Tajik-Cyrillic in Tajikistan. Despite the significant simila…
Improving Image Classification with Location Context
Kevin Tang, Manohar Paluri, Li Fei-Fei +2
With the widespread availability of cellphones and cameras that have GPS capabilities, it is common for images being uploaded to the Internet today to have GPS coordinates associat…
Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
Ayushi Pandey, Pamir Gogoi, Kevin Tang
Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-…
Prompt-to-Product: Generative Assembly via Bimanual Manipulation
Ruixuan Liu, Philip Huang, Ava Pun +8
Creating assembly products demands significant manual effort and expert knowledge in 1) designing the assembly and 2) constructing the product. This paper introduces Prompt-to-Prod…
Disambiguation of morpho-syntactic features of African American English -- the case of habitual be
Harrison Santiago, Joshua Martin, Sarah Moeller +1
Recent research has highlighted that natural language processing (NLP) systems exhibit a bias against African American speakers. The bias errors are often caused by poor representa…
Automatic Speech Recognition Biases in Newcastle English: an Error Analysis
Dana Serditova, Kevin Tang, Jochen Steffens
Automatic Speech Recognition (ASR) systems struggle with regional dialects due to biased training which favours mainstream varieties. While previous research has identified racial,…
Automatic Speech Recognition of African American English: Lexical and Contextual Effects
Hamid Mojarad, Kevin Tang
Automatic Speech Recognition (ASR) models often struggle with the phonetic, phonological, and morphosyntactic features found in African American English (AAE). This study focuses o…
GreedLlama: Performance of Financial Value-Aligned Large Language Models in Moral Reasoning
Jeffy Yu, Maximilian Huber, Kevin Tang
This paper investigates the ethical implications of aligning Large Language Models (LLMs) with financial optimization, through the case study of GreedLlama, a model fine-tuned to p…
Probing Character-level Transformers for the Spanish L-shaped Morphome
Akhilesh Kakolu Ramarao, Kevin Tang, Wiebke Petersen +1
When a transformer learns an irregular morphological pattern, what has it learned? Our test case is the Spanish \emph{L-shaped morphome}, a complex irregular pattern in which the v…
Transformers over-extend what humans underlearn: the case of Spanish L-shaped morphome
Akhilesh Kakolu Ramarao, Kevin Tang, Dinah Baer-Henney
The cognitive reality of irregular morphological patterns has been debated for decades: do speakers extend them to novel forms, or are they lexical artifacts? A neural network trai…
Frequency matters: Modeling irregular morphological patterns in Spanish with Transformers
Akhilesh Kakolu Ramarao, Kevin Tang, Dinah Baer-Henney
Over the past decade, various studies have addressed how speakers solve the so-called `The Paradigm Cell Filling Problem' (PCFP) \citep{ackerman2009parts} across different language…
CharacterFlywheel: Scaling Iterative Improvement of Engaging and Steerable LLMs in Production
Yixin Nie, Lin Guan, Zhongyao Ma +19
This report presents CharacterFlywheel, an iterative flywheel process for improving large language models (LLMs) in production social chat applications across Instagram, WhatsApp,…
Character-aware Transformers Learn an Irregular Morphological Pattern Yet None Generalize Like Humans
Akhilesh Kakolu Ramarao, Kevin Tang, Dinah Baer-Henney
Whether neural networks can serve as cognitive models of morphological learning remains an open question. Recent work has shown that encoder-decoder models can acquire irregular pa…
SN 2017cfd: A Normal Type Ia Supernova Discovered Very Young
Xuhui Han, WeiKang Zheng, Benjamin E. Stahl +22
The Type~Ia supernova (SN~Ia) 2017cfd in IC~0511 (redshift z = 0.01209+- 0.00016$) was discovered by the Lick Observatory Supernova Search 1.6+-0.7 d after the fitted first-light t…
Playing with Voices: Tabletop Role-Playing Game Recordings as a Diarization Challenge
Lian Remme, Kevin Tang
This paper provides a proof of concept that audio of tabletop role-playing games (TTRPG) could serve as a challenge for diarization systems. TTRPGs are carried out mostly by conver…
Lick Observatory Supernova Search Follow-Up Program: Photometry Data Release of 93 Type Ia Supernovae
Benjamin E. Stahl, WeiKang Zheng, Thomas de Jaeger +52
We present BVRI and unfiltered light curves of 93 Type Ia supernovae (SNe Ia) from the Lick Observatory Supernova Search (LOSS) follow-up program conducted between 2005 and 2018. O…
VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point Correspondence
Yuchao Gu, Yipin Zhou, Bichen Wu +7
Current diffusion-based video editing primarily focuses on structure-preserved editing by utilizing various dense correspondences to ensure temporal consistency and motion alignmen…
Mapping Subnational Vulnerability to Inadequate Micronutrient Intake using a Bayesian Small Area Estimation Framework
Sahoko Ishida, Mohammed Osman, Ziyao Cui +8
Inadequate dietary micronutrient intake is a significant risk factor for deficiency and remains a major global health challenge. Nutrition programmes and interventions are most eff…
Correlation between prosody and pragmatics: A case study of the discourse marker hÄlÄ `now' in Persian
Soleiman Ghaderi, Moloud Asakereh, Kevin Tang
The paper investigates how the Persian discourse marker hālä ('now') functions in conversation, analyzing its pragmatic roles and the acoustic cues—duration and intensity—that sign…
Probing for Phonology in Self-Supervised Speech Representations: A Case Study on Accent Perception
Nitin Venkateswaran, Kevin Tang, Ratree Wayland
Traditional models of accent perception underestimate the role of gradient variations in phonological features which listeners rely upon for their accent judgments. We investigate…