Skeleton-of-Thought: Large Language Models Can Do Parallel Decoding
Xuefei Ning, Zinan Lin, Zixuan Zhou +2 authors
"Skeleton-of-Thought" (SoT) method reduces generation latency and potentially improves answer quality by guiding LLMs to generate a skeleton first, followed by parallel content completion.