My Publications

Here’s a list of my conference and journal publications, spanning my work on trustworthy AI, LLMs, NLP, and reliable machine learning systems.

2026

ConceptGuard: Benchmarking Context-Sensitive Unlearning in Large Language Models

Sahil Kale, Ian Harris

  • Enables LLM unlearning to be explored and gauged at the level of concepts, instead of sparse facts, with evaluation being intent-sensitive to maximize contextual separation and promote safer behavior

Submitted to NeurIPS E&D 2026, Sydney, Australia

Read the paper

Local Information Access in Marathi: Evaluating LLM-Native Web Retrieval in a Low-Resource Environment

Sahil Kale

  • Identifies challenges in AI-backed information retrieval for under-represented languages, isolating failure modes in translated global queries and natively crafted local information needs.

ACM SIGIR Conference 2026, Melbourne, Australia

Read the paper

Future Confidence Distillation in Large Language Models

Sahil Kale

  • Investigates how language models can improve their confidence estimation by distilling information about future states post answer generation.
  • Shows how distilled predictors recover calibration improvement achieved by post-solution confidence, remain highly sample efficient, and transfer across domains

Submitted to AAAI 2027, Montreal, Canada

Read the paper

Designing Policy with Last-Mile Stakeholders: Connecting Ground-Level Farmer Insights to Indian Agrarian Policymakers with NLP

Kasturi Pathak, Sahil Kale

  • Explores how NLP and voice-based AI can connect ground-level farmer insights with policymakers, enabling last-mile perspectives to inform agricultural policy design and formulation.

ACM DIS Conference 2026, Singapore

Read the paper

KnowRL: Teaching Language Models to Know What They Know

Sahil Kale, Devendra Singh Dhami

  • Presents a framework that strengthens a model’s internal understanding of its own feasibility boundaries using only a small seed set and no external supervision, achieving gains of up to 28% in accuracy and 12% in F1.

Submitted to NeurIPS 2026, Sydney, Australia

Read the paper

Lie to Me: Knowledge Graphs for Robust Hallucination Self-Detection in LLMs

Sahil Kale, Antonio Luca Alfeo

  • Demonstrates that structuring LLM outputs as knowledge graphs significantly improves hallucination self-detection, achieving up to 16% higher accuracy and 20% better F1-score over existing methods.

ICPRAM Conference 2026, Marbella, Spain

Read the paper

i-Check: An Idempotence-Driven Optimisation Framework for AI Agents in Enterprise Workflows

Sahil Kale, Yash Nikam, Vijaykant Nadadur

  • Introduces input constraints called idempotence-driving constraints for agents to achieve enhanced repeatability and consistency up to 90% across responses, while reducing costs by 37% in enterprise workflows.

ICAART Conference 2026, Marbella, Spain

Read the paper

Mirage of Mastery: Memorization Tricks LLMs into Artificially Inflated Self-Knowledge

Sahil Kale

  • Shows that LLMs can draw confidence from memorized solutions to infer artificially inflated self-knowledge about their reasoning ability, resulting in over 45% inconsistency in feasibility assessments.

AAAI IASEAI Conference 2026, Paris, France

Read the paper

Geospatial modeling study assessing population level accessibility to medical college hospitals in India

Harsh Thakkar, Chaitanya Reddy, Varun Raj Passi, Aamir Miyajiwala, Sahil Kale, Ankit Raj, Siddhesh Zadey

  • Uses geocoding and spatial modeling to assess accessibility to medical college hospitals across India.
  • Presents the density of MCHs, median travel times, and Access Population Coverage (APC) across 36 states and 735 districts, revealing significant disparities in access

Discover Public Health, Vol 23, 2026

Read the paper

2025

Look It Up: Analysing Internal Web Search Capabilities of Modern LLMs

Sahil Kale

  • Investigates whether modern LLMs understand when external web search is necessary and what information they should search for.
  • Examines the internal web-search behavior of language models as part of their broader ability to recognize and address knowledge gaps.

arXiv preprint

Read the paper

A secure and imperceptible communication system for sharing co-ordinate data

Ranjeet Bidwe, Sahil Kale, Gautam Khaire, Jay Patankar, Deepak Mane, Suraj Sawant

  • Combines AES encryption with hash-driven multi-image steganography to enable secure and imperceptible transmission of high-volume military coordinate data.
  • Evaluates the approach as a practical and computationally efficient solution for sensitive communication channels.

Scientific Reports, Vol. 15, 2025

Read the paper

TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs

Sahil Kale, Vijaykant Nadadur

  • Proposes TeXpert, a benchmark with natural-language prompts for generating LaTeX components of scientific documents across multiple difficulty levels.
  • Shows that LLM performance remains poor on complex LaTeX generation despite strong performance on standard benchmarks.

Scholarly Document Processing Workshop @ ACL 2025, Vienna, Austria

Read the paper

Line of Duty: Evaluating LLM Self-Knowledge via Consistency in Feasibility Boundaries

Sahil Kale, Vijaykant Nadadur

  • Introduces a methodology for obtaining intrinsic insights into LLM self-knowledge through consistency in self-defined feasibility boundaries.
  • Found that even frontier models such as GPT-4o and Mistral Large are uncertain about their capabilities more than 80% of the time.

TrustAI Workshop @NAACL 2025, Albuquerque, USA

Read the paper

2024

FAQ-Gen: An automated system to generate domain-specific FAQs to aid content comprehension

Sahil Kale, Gautam Khaire, Jay Patankar

  • Proposes an FAQ generation system using custom-built and fine-tuned text-to-text transformation models with self-curated algorithms for cognitive ranking of question-answer pairs.

Journal of Computer-Assisted Linguistic Research, Universitat Politècnica de València, Vol. 8, 2024

Read the paper

2023

A Modern Approach to Electoral Delimitation using the Quadtree Data Structure

Sahil Kale, Gautam Khaire, Jay Patankar, Pujashree Vidap

  • Proposes a novel system using the quadtree data structure to automatically demarcate constituency boundaries while satisfying the requirements for fair delimitation established by the US Supreme Court.

IEEE ICCCEE Conference 2023, Pune, India

Read the paper