الباحثون

Yu Huang

المنشورات 7

نسخة أولية وصول مفتوح

UltraText Bench: A Comprehensive Bilingual Benchmark for Evaluating Visual Text Rendering in Image Generation

Deyuan Liu, Yihao Hu, Jingxuan Zhang وآخرون · 2026

Dense visual text requires image generators to reproduce long strings across multiple regions with correct placement and legibility. As short-string rendering improves, evaluation must test sustained performance across more demanding scenes. We introduce UltraText Bench, a bilingual benchmark for prompt-only generation …

نسخة أولية وصول مفتوح

VTR-Bench: A Systematic Benchmark for Evaluating Visual Text Rendering in Video Generation

Yu Huang, Jungang Li, Zhiyuan Wang وآخرون · 2026

Recent video generation models can produce highly realistic videos from natural language instructions, with visual quality approaching cinematic standards. Existing evaluation benchmarks, however, predominantly assess visual quality, aesthetic appeal and physical plausibility, while paying limited attention to text, an …

نسخة أولية وصول مفتوح

Provable Test-Time Scaling for Beam Search in LLM Reasoning

Qijia He, Yu Huang, Yu An Cheng وآخرون · 2026

Beam-search-based test-time methods provide an effective way to improve large language model (LLM) performance on long-horizon generation by pruning invalid reasoning paths early, leading to significantly improved reasoning efficiency and more favorable test-time cost scaling. Despite strong empirical success, the theo …

نسخة أولية وصول مفتوح

TCSAlgBench: Benchmarking Automated Proving for Research-Level Theoretical Computer Science

Chutong Yang, Xiyuan Zhang, Yu Huang وآخرون · 2026

Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretical computer science (TCS) connects algorithm design to explicit guarantees and fundamental limits, providing a setting for evaluating whether models can justify co …

نسخة أولية وصول مفتوح

InterTab: Interleaved Visual-Structure Alignment for Multi-Modal Table Reasoning

Hanqian Li, Sirui Huang, Chen Ling وآخرون · 2026

Table images preserve structural information that are often lost in text serialization, and reasoning over them requires locating relevant rows, columns, and cells step by step. Current multimodal large language models (MLLMs) encode the whole image once before reasoning, so they cannot pick up row-, column-, and cell- …

نسخة أولية وصول مفتوح

HyperCLIP++: Fine-tuning CLIP forOpen-vocabulary Semantic Segmentation in Hyperbolic Space

Zelin Peng, Zhengqin Xu, Changsong Wen وآخرون · 2026

CLIP, a foundational vision-language model, has emerged as a powerful tool for open-vocabulary semantic segmentation. While freezing CLIP's text encoder is known to preserve its generalization capability, recent studies show that fine-tuning both CLIP's text and image encoders jointly significantly enhances segmentatio …

المؤلفون المشاركون