Skip to main content
Commands, package names, and image names on this page come from the open-source project that Mibyan Desktop is built on, and can differ from the Mibyan Desktop installer. For the supported Mibyan install and update path, see Install and update.
Write ML papers for NeurIPS/ICML/ICLR: design→submit.

Skill metadata

Reference: full SKILL.md

The following is the complete skill definition that Mibyan loads when this skill is triggered. This is what the agent sees as instructions when the skill is active.

Research Paper Writing Pipeline

End-to-end pipeline for producing publication-ready ML/AI research papers targeting NeurIPS, ICML, ICLR, ACL, AAAI, and COLM. This skill covers the full research lifecycle: experiment design, execution, monitoring, analysis, paper writing, review, revision, and submission. This is not a linear pipeline — it is an iterative loop. Results trigger new experiments. Reviews trigger new analysis. The agent must handle these feedback loops.

When To Use This Skill

Use this skill when:
  • Starting a new research paper from an existing codebase or idea
  • Designing and running experiments to support paper claims
  • Writing or revising any section of a research paper
  • Preparing for submission to a specific conference or workshop
  • Responding to reviews with additional experiments or revisions
  • Converting a paper between conference formats
  • Writing non-empirical papers — theory, survey, benchmark, or position papers (see Paper Types Beyond Empirical ML)
  • Designing human evaluations for NLP, HCI, or alignment research
  • Preparing post-acceptance deliverables — posters, talks, code releases

Core Philosophy

  1. Be proactive. Deliver complete drafts, not questions. Scientists are busy — produce something concrete they can react to, then iterate.
  2. Never hallucinate citations. AI-generated citations have ~40% error rate. Always fetch programmatically. Mark unverifiable citations as [CITATION NEEDED].
  3. Paper is a story, not a collection of experiments. Every paper needs one clear contribution stated in a single sentence. If you can’t do that, the paper isn’t ready.
  4. Experiments serve claims. Every experiment must explicitly state which claim it supports. Never run experiments that don’t connect to the paper’s narrative.
  5. Commit early, commit often. Every completed experiment batch, every paper draft update — commit with descriptive messages. Git log is the experiment history.

Proactivity and Collaboration

Default: Be proactive. Draft first, ask with the draft. Block for input only when: target venue unclear, multiple contradictory framings, results seem incomplete, explicit request to review first.

Phase 0: Project Setup

Goal: Establish the workspace, understand existing work, identify the contribution.

Step 0.1: Explore the Repository

Look for:
  • README.md — project overview and claims
  • results/, outputs/, experiments/ — existing findings
  • configs/ — experimental settings
  • .bib files — existing citations
  • Draft documents or notes

Step 0.2: Organize the Workspace

Establish a consistent workspace structure:

Step 0.3: Set Up Version Control

Git discipline: Every completed experiment batch gets committed with a descriptive message. Example:

Step 0.4: Identify the Contribution

Before writing anything, articulate:
  • The What: What is the single thing this paper contributes?
  • The Why: What evidence supports it?
  • The So What: Why should readers care?
Propose to the scientist: “Based on my understanding, the main contribution is: [one sentence]. The key results show [Y]. Is this the framing you want?”

Step 0.5: Create a TODO List

Use the todo tool to create a structured project plan:
Update this throughout the project. It serves as the persistent state across sessions.

Step 0.6: Estimate Compute Budget

Before running experiments, estimate total cost and time:
Track actual spend as experiments run:
When budget is tight: Run pilot experiments (1-2 seeds, subset of tasks) before committing to full sweeps. Use cheaper models for debugging pipelines, then switch to target models for final runs.

Step 0.7: Multi-Author Coordination

Most papers have 3-10 authors. Establish workflows early: Section ownership: Assign each section to one primary author. Others comment but don’t edit directly. Prevents merge conflicts and style inconsistency.
LaTeX conventions to agree on early:
  • \method{} macro for consistent method naming
  • Citation style: \citet{} vs \citep{} usage
  • Math notation: lowercase bold for vectors, uppercase bold for matrices, etc.
  • British vs American spelling

Phase 1: Literature Review

Goal: Find related work, identify baselines, gather citations.

Step 1.1: Identify Seed Papers

Start from papers already referenced in the codebase:
Load the arxiv skill for structured paper discovery: skill_view("arxiv"). It provides arXiv REST API search, Semantic Scholar citation graphs, author profiles, and BibTeX generation. Use web_search for broad discovery, web_extract for fetching specific papers:
Additional search queries to try:
Recommended: Install Exa MCP for real-time academic search:

Step 1.2b: Deepen the Search (Breadth-First, Then Depth)

A flat search (one round of queries) typically misses important related work. Use an iterative breadth-then-depth pattern inspired by deep research pipelines:
When to stop: If a round returns >80% papers already in your collection, the search is saturated. Typically 2-3 rounds suffice. For survey papers, expect 4-5 rounds. For agent-based workflows: Delegate each round’s queries in parallel via delegate_task. Collect results, deduplicate, then generate the next round’s queries from the combined learnings.

Step 1.3: Verify Every Citation

NEVER generate BibTeX from memory. ALWAYS fetch programmatically. For each citation, follow the mandatory 5-step process:
If you cannot verify a citation:
Always tell the scientist: “I’ve marked [X] citations as placeholders that need verification.” See references/citation-workflow.md for complete API documentation and the full CitationManager class. Group papers by methodology, not paper-by-paper: Good: “One line of work uses X’s assumption [refs] whereas we use Y’s assumption because…” Bad: “Smith et al. introduced X. Jones et al. introduced Y. We combine both.”

Phase 2: Experiment Design

Goal: Design experiments that directly support paper claims. Every experiment must answer a specific question.

Step 2.1: Map Claims to Experiments

Create an explicit mapping: Rule: If an experiment doesn’t map to a claim, don’t run it.

Step 2.2: Design Baselines

Strong baselines are what separates accepted papers from rejected ones. Reviewers will ask: “Did they compare against X?” Standard baseline categories:
  • Naive baseline: Simplest possible approach
  • Strong baseline: Best known existing method
  • Ablation baselines: Your method minus one component
  • Compute-matched baselines: Same compute budget, different allocation

Step 2.3: Define Evaluation Protocol

Before running anything, specify:
  • Metrics: What you’re measuring, direction symbols (higher/lower better)
  • Aggregation: How results are combined across runs/tasks
  • Statistical tests: What tests will establish significance
  • Sample sizes: How many runs/problems/tasks

Step 2.4: Write Experiment Scripts

Follow these patterns from successful research pipelines: Incremental saving — save results after each step for crash recovery:
Artifact preservation — save all intermediate outputs:
Separation of concerns — keep generation, evaluation, and visualization separate:
See references/experiment-patterns.md for complete design patterns, cron monitoring, and error recovery.

Step 2.5: Design Human Evaluation (If Applicable)

Many NLP, HCI, and alignment papers require human evaluation as primary or complementary evidence. Design this before running automated experiments — human eval often has longer lead times (IRB approval, annotator recruitment). When human evaluation is needed:
  • Automated metrics don’t capture what you care about (fluency, helpfulness, safety)
  • Your contribution is about human-facing qualities (readability, preference, trust)
  • Reviewers at NLP venues (ACL, EMNLP) expect it for generation tasks
Key design decisions: Annotation guideline checklist:
Reporting requirements (reviewers check all of these):
  • Number of annotators and their qualifications
  • Inter-annotator agreement with specific metric and value
  • Compensation details (amount, estimated hourly rate)
  • Annotation interface description or screenshot (appendix)
  • Total annotation time
See references/human-evaluation.md for complete guide including statistical tests for human eval data, crowdsourcing quality control patterns, and IRB guidance.

Phase 3: Experiment Execution & Monitoring

Goal: Run experiments reliably, monitor progress, recover from failures.

Step 3.1: Launch Experiments

Use nohup for long-running experiments:
Parallel execution: Run independent experiments simultaneously, but be aware of API rate limits. 4+ concurrent experiments on the same API will slow each down.

Step 3.2: Set Up Monitoring (Cron Pattern)

For long-running experiments, set up periodic status checks. The cron prompt should follow this template:
Silent mode: If nothing has changed since the last check, respond with [SILENT] to suppress notification to the user. Only report when there’s news.

Step 3.3: Handle Failures

Common failure modes and recovery: Key: Scripts should always check for existing results and skip completed work. This makes re-runs safe and efficient.

Step 3.4: Commit Completed Results

After each experiment batch completes:

Step 3.5: Maintain an Experiment Journal

Git commits track what happened, but not the exploration tree — the decisions about what to try next based on what you learned. Maintain a structured experiment journal that captures this tree:
Why a journal, not just git? Git tracks file changes. The journal tracks the reasoning: why you tried X, what you learned, and what that implies for the next experiment. When writing the paper, this tree is invaluable for the Methods section (“we observed X, which motivated Y”) and for honest failure reporting. Selecting the best path: When the journal shows a branching tree (exp_001 → exp_002a, exp_002b, exp_003), identify the path that best supports the paper’s claims. Document dead-end branches in the appendix as ablations or negative results. Snapshot code per experiment: Copy the experiment script after each run:
This enables exact reproduction even after subsequent code changes.

Phase 4: Result Analysis

Goal: Extract findings, compute statistics, identify the story.

Step 4.1: Aggregate Results

Write analysis scripts that:
  1. Load all result files from a batch
  2. Compute per-task and aggregate metrics
  3. Generate summary tables

Step 4.2: Statistical Significance

Always compute:
  • Error bars: Standard deviation or standard error, specify which
  • Confidence intervals: 95% CI for key results
  • Pairwise tests: McNemar’s test for comparing two methods
  • Effect sizes: Cohen’s d or h for practical significance
See references/experiment-patterns.md for complete implementations of McNemar’s test, bootstrapped CIs, and Cohen’s h.

Step 4.3: Identify the Story

After analysis, explicitly answer:
  1. What is the main finding? State it in one sentence.
  2. What surprised you? Unexpected results often make the best papers.
  3. What failed? Failed experiments can be the most informative. Honest reporting of failures strengthens the paper.
  4. What follow-up experiments are needed? Results often raise new questions.

Handling Negative or Null Results

When your hypothesis was wrong or results are inconclusive, you have three options: How to write a negative results paper:
  • Lead with what the community believes and why it matters to test it
  • Describe your rigorous methodology (must be airtight — reviewers will scrutinize harder)
  • Present the null result clearly with statistical evidence
  • Analyze why the expected result didn’t materialize
  • Discuss implications for the field
Venues that explicitly welcome negative results: NeurIPS (Datasets & Benchmarks track), TMLR, ML Reproducibility Challenge, workshops at major conferences. Some workshops specifically call for negative results.

Step 4.4: Create Figures and Tables

Figures:
  • Use vector graphics (PDF) for all plots: plt.savefig('fig.pdf')
  • Colorblind-safe palettes (Okabe-Ito or Paul Tol)
  • Self-contained captions — reader should understand without main text
  • No title inside figure — the caption serves this function
Tables:
  • Use booktabs LaTeX package
  • Bold best value per metric
  • Include direction symbols (higher/lower better)
  • Consistent decimal precision

Step 4.5: Decide: More Experiments or Write?

Step 4.6: Write the Experiment Log (Bridge to Writeup)

Before moving to paper writing, create a structured experiment log that bridges results to prose. This is the single most important connective tissue between experiments and the writeup — without it, the writing agent has to re-derive the story from raw result files. Create experiment_log.md with the following structure:
Why this matters: When drafting, the agent (or a delegated sub-agent) can load experiment_log.md alongside the LaTeX template and produce a first draft grounded in actual results. Without this bridge, the writing agent must parse raw JSON/CSV files and infer the story — a common source of hallucinated or misreported numbers. Git discipline: Commit this log alongside the results it describes.

Iterative Refinement: Strategy Selection

Any output in this pipeline — paper drafts, experiment scripts, analysis — can be iteratively refined. The autoreason research provides empirical evidence for when each refinement strategy works and when it fails. Use this section to choose the right approach.

Quick Decision Table

The Generation-Evaluation Gap

Core insight: Autoreason’s value depends on the gap between a model’s generation capability and its self-evaluation capability.
This gap is structural, not temporary. As costs drop, today’s frontier becomes tomorrow’s mid-tier. The sweet spot moves but never disappears.

Autoreason Loop (Summary)

Each pass produces three candidates from fresh, isolated agents:
  1. Critic → finds problems in incumbent A (no fixes)
  2. Author B → revises A based on critique
  3. Synthesizer → merges A and B (randomized labels)
  4. Judge Panel → 3 blind CoT judges rank A, B, AB via Borda count
  5. Convergence → A wins k=2 consecutive passes → done
Key parameters:
  • k=2 convergence (k=1 premature, k=3 too expensive, no quality gain)
  • CoT judges always (3x faster convergence)
  • Temperature 0.8 authors, 0.3 judges
  • Conservative tiebreak: incumbent wins ties
  • Every role is a fresh agent with no shared context

Applying to Paper Drafts

When refining the paper itself through autoreason:
  • Provide ground truth to the critic: actual experimental data, result JSONs, statistical outputs. Without this, models hallucinate fabricated ablation studies and fake confidence intervals.
  • Use 3 working judges minimum: A broken judge parser doesn’t add noise — it prevents equilibrium entirely.
  • Scope constrain the revision: “Address these specific weaknesses” not “improve the paper.”

Failure Modes

See references/autoreason-methodology.md for complete prompts, Borda scoring details, model selection guide, scope constraint design patterns, and compute budget reference.

Phase 5: Paper Drafting

The complete drafting procedure (section-by-section order, LaTeX scaffolding, figure/table conventions, abstract and intro formulas, related-work positioning) lives in references/phase5-paper-drafting.md — load it with read_file when you reach this phase. Pair it with references/writing-guide.md for prose-level style rules.

Phase 6: Self-Review & Revision

Goal: Simulate the review process before submission. Catch weaknesses early.

Step 6.1: Simulate Reviews (Ensemble Pattern)

Generate reviews from multiple perspectives. The key insight from automated research pipelines (notably SakanaAI’s AI-Scientist): ensemble reviewing with a meta-reviewer produces far more calibrated feedback than a single review pass. Step 1: Generate N independent reviews (N=3-5) Use different models or temperature settings. Each reviewer sees only the paper, not other reviews. Default to negative bias — LLMs have well-documented positivity bias in evaluation.
Step 2: Meta-review (Area Chair aggregation) Feed all N reviews to a meta-reviewer:
Step 3: Reflection loop (optional, 2-3 rounds) Each reviewer can refine their review after seeing the meta-review. Use an early termination sentinel: if the reviewer responds “I am done” (no changes), stop iterating. Model selection for reviewing: Reviewing is best done with the strongest available model, even if you wrote the paper with a cheaper one. The reviewer model should be chosen independently from the writing model. Few-shot calibration: If available, include 1-2 real published reviews from the target venue as examples. This dramatically improves score calibration. See references/reviewer-guidelines.md for example reviews.

Step 6.1b: Visual Review Pass (VLM)

Text-only review misses an entire class of problems: figure quality, layout issues, visual consistency. If you have access to a vision-capable model, run a separate visual review on the compiled PDF:
This catches problems that text-based review cannot: a plot with illegible axis labels, a figure placed 3 pages from its first reference, inconsistent color palettes between Figure 2 and Figure 5, or a table that’s clearly wider than the column width.

Step 6.1c: Claim Verification Pass

After simulated reviews, run a separate verification pass. This catches factual errors that reviewers might miss:
For agent-based workflows: delegate verification to a fresh sub-agent that receives only the paper text and the raw result files. The fresh context prevents confirmation bias — the verifier doesn’t “remember” what the results were supposed to be.

Step 6.2: Prioritize Feedback

After collecting reviews, categorize:

Step 6.3: Revision Cycle

For each critical/high issue:
  1. Identify the specific section(s) affected
  2. Draft the fix
  3. Verify the fix doesn’t break other claims
  4. Update the paper
  5. Re-check against the reviewer’s concern

Step 6.4: Rebuttal Writing

When responding to actual reviews (post-submission), rebuttals are a distinct skill from revision: Format: Point-by-point. For each reviewer concern:
Rules:
  • Address every concern — reviewers notice if you skip one
  • Lead with the strongest responses
  • Be concise and direct — reviewers read dozens of rebuttals
  • Include new results if you ran experiments during the rebuttal period
  • Never be defensive or dismissive, even of weak criticisms
  • Use latexdiff to generate a marked-up PDF showing changes (see Professional LaTeX Tooling section)
  • Thank reviewers for specific, actionable feedback (not generic praise)
What NOT to do: “We respectfully disagree” without evidence. “This is out of scope” without explanation. Ignoring a weakness by only responding to strengths.

Step 6.5: Paper Evolution Tracking

Save snapshots at key milestones:

Phase 7: Submission Preparation

Goal: Final checks, formatting, and submission.

Step 7.1: Conference Checklist

Every venue has mandatory checklists. Complete them carefully — incomplete checklists can result in desk rejection. See references/checklists.md for:
  • NeurIPS 16-item paper checklist
  • ICML broader impact + reproducibility
  • ICLR LLM disclosure policy
  • ACL mandatory limitations section
  • Universal pre-submission checklist

Step 7.2: Anonymization Checklist

Double-blind review means reviewers cannot know who wrote the paper. Check ALL of these:
Common mistakes: Git commit messages visible in supplementary code, watermarked figures from institutional tools, acknowledgments left in from a previous draft, arXiv preprint posted before anonymity period.

Step 7.3: Formatting Verification

Step 7.4: Pre-Compilation Validation

Run these automated checks before attempting pdflatex. Catching errors here is faster than debugging compiler output.
Fix any warnings before proceeding. For agent-based workflows: feed chktex output back to the agent with instructions to make minimal fixes.

Step 7.5: Final Compilation

If compilation fails: Parse the .log file for the first error. Common fixes:
  • “Undefined control sequence” → missing package or typo in command name
  • “Missing $ inserted” → math symbol outside math mode
  • “File not found” → wrong figure path or missing .sty file
  • “Citation undefined” → .bib entry missing or bibtex not run

Step 7.6: Conference-Specific Requirements

Step 7.7: Conference Resubmission & Format Conversion

When converting between venues, never copy LaTeX preambles between templates:
When cutting pages: move proofs to appendix, condense related work, combine tables, use subfigures. When expanding: add ablations, expand limitations, include additional baselines, add qualitative examples. After rejection: Address reviewer concerns in the new version, but don’t include a “changes” section or reference the previous submission (blind review).

Step 7.8: Camera-Ready Preparation (Post-Acceptance)

After acceptance, prepare the camera-ready version:

Step 7.9: arXiv & Preprint Strategy

Posting to arXiv is standard practice in ML but has important timing and anonymity considerations. Timing decision tree: arXiv category selection (ML/AI papers): List primary + 1-2 cross-listed categories. More categories = more visibility, but only cross-list where genuinely relevant. Versioning strategy:
  • v1: Initial submission (matches conference submission)
  • v2: Post-acceptance with camera-ready corrections (add “accepted at [Venue]” to abstract)
  • Don’t post v2 during the review period with changes that clearly respond to reviewer feedback

Step 7.10: Research Code Packaging

Releasing clean, runnable code significantly increases citations and reviewer trust. Package code alongside the camera-ready submission. Repository structure:
README template for research code:
Pre-release checklist:
Anonymous code for submission (before acceptance):

Phase 8: Post-Acceptance Deliverables

Goal: Maximize the impact of your accepted paper through presentation materials and community engagement.

Step 8.1: Conference Poster

Most conferences require a poster session. Poster design principles: Tools: LaTeX (beamerposter package), PowerPoint/Keynote, Figma, Canva. Production: Order 2+ weeks before the conference. Fabric posters are lighter for travel. Many conferences now support virtual/digital posters too.

Step 8.2: Conference Talk / Spotlight

If awarded an oral or spotlight presentation: Slide design rules:
  • One idea per slide
  • Minimize text — speak the details, don’t project them
  • Animate key figures to build understanding step-by-step
  • Include a “takeaway” slide at the end (single sentence contribution)
  • Prepare backup slides for anticipated questions

Step 8.3: Blog Post / Social Media

An accessible summary significantly increases impact:
  • Twitter/X thread: 5-8 tweets. Lead with the result, not the method. Include Figure 1 and key result figure.
  • Blog post: 800-1500 words. Written for ML practitioners, not reviewers. Skip formalism, emphasize intuition and practical implications.
  • Project page: HTML page with abstract, figures, demo, code link, BibTeX. Use GitHub Pages.
Timing: Post within 1-2 days of paper appearing on proceedings or arXiv camera-ready.

Workshop & Short Papers

Workshop papers and short papers (e.g., ACL short papers, Findings papers) follow the same pipeline but with different constraints and expectations.

Workshop Papers

When to target a workshop:
  • Early-stage idea you want feedback on before a full paper
  • Negative result that doesn’t justify 8+ pages
  • Position piece or opinion on a timely topic
  • Replication study or reproducibility report

ACL Short Papers & Findings

ACL venues have distinct submission types: Short paper strategy: Pick ONE claim and support it thoroughly. Don’t try to compress a long paper into 4 pages — write a different, more focused paper.

Paper Types Beyond Empirical ML

The main pipeline above targets empirical ML papers. Other paper types require different structures and evidence standards. See references/paper-types.md for detailed guidance on each type.

Theory Papers

Structure: Introduction → Preliminaries (definitions, notation) → Main Results (theorems) → Proof Sketches → Discussion → Full Proofs (appendix) Key differences from empirical papers:
  • Contribution is a theorem, bound, or impossibility result — not experimental numbers
  • Methods section replaced by “Preliminaries” and “Main Results”
  • Proofs are the evidence, not experiments (though empirical validation of theory is welcome)
  • Proof sketches in main text, full proofs in appendix is standard practice
  • Experimental section is optional but strengthens the paper if it validates theoretical predictions
Proof writing principles:
  • State theorems formally with all assumptions explicit
  • Provide intuition before formal proof (“The key insight is…”)
  • Proof sketches should convey the main idea in 0.5-1 page
  • Use \begin{proof}...\end{proof} environments
  • Number assumptions and reference them in theorems: “Under Assumptions 1-3, …”

Survey / Tutorial Papers

Structure: Introduction → Taxonomy / Organization → Detailed Coverage → Open Problems → Conclusion Key differences:
  • Contribution is the organization, synthesis, and identification of open problems — not new methods
  • Must be comprehensive within scope (reviewers will check for missing references)
  • Requires a clear taxonomy or organizational framework
  • Value comes from connections between works that individual papers don’t make
  • Best venues: TMLR (survey track), JMLR, Foundations and Trends in ML, ACM Computing Surveys

Benchmark Papers

Structure: Introduction → Task Definition → Dataset Construction → Baseline Evaluation → Analysis → Intended Use & Limitations Key differences:
  • Contribution is the benchmark itself — it must fill a genuine evaluation gap
  • Dataset documentation is mandatory, not optional (see Datasheets, Step 5.11)
  • Must demonstrate the benchmark is challenging (baselines don’t saturate it)
  • Must demonstrate the benchmark measures what you claim it measures (construct validity)
  • Best venues: NeurIPS Datasets & Benchmarks track, ACL (resource papers), LREC-COLING

Position Papers

Structure: Introduction → Background → Thesis / Argument → Supporting Evidence → Counterarguments → Implications Key differences:
  • Contribution is an argument, not a result
  • Must engage seriously with counterarguments
  • Evidence can be empirical, theoretical, or logical analysis
  • Best venues: ICML (position track), workshops, TMLR

Mibyan Integration

This skill is designed for the Mibyan agent. It uses Mibyan tools, delegation, scheduling, and memory for the full research lifecycle. Compose this skill with other Mibyan skills for specific phases: This skill supersedes ml-paper-writing — it contains all of ml-paper-writing’s content plus the full experiment/analysis pipeline and autoreason methodology.

Mibyan Tools Reference

Tool Usage Patterns

Experiment monitoring (most common):
Parallel section drafting (using delegation):
Each delegate runs as a fresh subagent with no shared context — provide all necessary information in the prompt. Collect outputs and integrate. Citation verification (using execute_code):

State Management with memory and todo

memory tool — persist key decisions (bounded: ~2200 chars for MEMORY.md):
Update memory after major decisions or phase transitions. This persists across sessions. todo tool — track granular progress:
Session startup protocol:

Cron Monitoring with cronjob

Use the cronjob tool to schedule periodic experiment checks:
[SILENT] protocol: When nothing has changed since the last check, respond with exactly [SILENT]. This suppresses notification delivery to the user. Only report when there are genuine changes worth knowing about. Deadline tracking:

Communication Patterns

When to notify the user (via your direct/final response, or a cron deliver: target for unattended runs):
  • Experiment batch completed (with results table)
  • Unexpected finding or failure requiring decision
  • Draft section ready for review
  • Deadline approaching with incomplete tasks
When NOT to notify:
  • Experiment still running, no new results → [SILENT]
  • Routine monitoring with no changes → [SILENT]
  • Intermediate steps that don’t need attention
Report format — always include structured data:

Decision Points Requiring Human Input

Use clarify for targeted questions when genuinely blocked: Do NOT ask about (be proactive, make a choice, flag it):
  • Word choice, section ordering
  • Which specific results to highlight
  • Citation completeness (draft with what you find, note gaps)

Reviewer Evaluation Criteria

Understanding what reviewers look for helps focus effort: Scoring (NeurIPS 6-point scale):
  • 6: Strong Accept — groundbreaking, flawless
  • 5: Accept — technically solid, high impact
  • 4: Borderline Accept — solid, limited evaluation
  • 3: Borderline Reject — weaknesses outweigh
  • 2: Reject — technical flaws
  • 1: Strong Reject — known results or ethics issues
See references/reviewer-guidelines.md for detailed guidelines, common concerns, and rebuttal strategies.

Common Issues and Solutions


Reference Documents

LaTeX Templates

Templates in templates/ for: NeurIPS 2025, ICML 2026, ICLR 2026, ACL, AAAI 2026, COLM 2025. See templates/README.md for compilation instructions.

Key External Sources

Writing Philosophy: APIs: Semantic Scholar | CrossRef | arXiv Venues: NeurIPS | ICML | ICLR | ACL