I'm Caleb β€” a full-stack engineer with a background in AI/ML. I'm currently a Lead Software Engineer at Open Government Products, where I build tech for public good.

I am Singaporean πŸ‡ΈπŸ‡¬, so ask me about local food. In my free time, I enjoy reading and playing chess β™ŸοΈ badly.

stuff I've worked on

cover image

Peek shows patients live, personalised pharmacy wait times powered by machine learning, and explains each stage of the medication preparation process. Piloted at Tan Tock Seng Hospital, it is used by a third of patients daily, with 93% rating it helpful.

cover image

Scribe is an AI transcription and summarisation tool that transforms conversations into structured notes, reducing documentation time by 60%. It is used by all public healthcare institutions and more than 100 social service agencies in Singapore.

cover image

Care360 is a patient management system for Medical Social Workers that supports planning and management of patients' psychosocial and/or financial care. It is used by all public healthcare institutions in Singapore.

cover image

Cleanlab's flagship SaaS product -- a no-code automatic data correction solution for fixing mislabeled data, removing out-of-distribution data, and ranking data by quality.

cover image

Cleanlab automatically finds errors in datasets and enables users to do machine learning/analytics with messy real-world data and labels.

cover image

Stanford's chatbot submission for the Alexa Prize competition, a university competition to advance conversational AI. It incorporates state-of-the-art methods in controllable neural generation and rich response generation modules that are informed by the hundreds of thousands of people that chatted with it throughout the year. In 2021, we came in 2nd and earned a $100K USD cash prize.

cover image

ConvoKit is a Python toolkit for extracting conversational features and analyzing social phenomena in conversations. It includes several large conversational datasets along with scripts exemplifying the use of the toolkit on these datasets.