We evaluate DeepCode on the PaperBench benchmark (released by OpenAI), a rigorous testbed requiring AI agents to independently reproduce 20 ICML 2024 papers from scratch. The benchmark comprises 8,316 ...
Abstract: Automated program repair (APR) aims to help developers improve software reliability by generating patches for buggy programs. Although many code language models (CLM) are developed and ...
Abstract: Joint exploration of intra/inter-video coding versatility in signal processing domain and signature/hash diversity in the information security domain has not been well investigated in the ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results