Measure whether a post-training change improved real enterprise work.
Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, production agent evaluation sets, and LLM-systems measurement for post-training loops.
Evaluation under uncertainty
- Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
- Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
- That is the measurement analogue of rigorous measures for complex, real-world enterprise tasks.
Agent evaluation sets
- Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium, including a multi-tenant conversational receptionist.
- That maps to inspecting multi-step traces for planning and tool use. It is not Writer Palmyra post-training or writing-product platform research.
- Ph.D. or equivalent demonstrated research experience is a posted requirement; this profile includes an Environmental Engineering PhD.
Proposed first contribution
For one post-training or agentic-workflow evaluation already in flight, define what performance on complex, real-world enterprise tasks means versus a score that is easy to move. Write a small failure taxonomy: metric movement without an enterprise-task quality change, slice-specific collapse on tool use or planning, an automatic judge that disagrees with intended behavior, coherence loss on a long-running task that still looks like a benchmark win. Stand up a small evaluation set with provenance on traces, compare simple baselines, attach uncertainty, and write a clear report before expanding SFT, RLHF, RLAIF, or DPO work. This is a proposed measurement approach, not a claim of prior Writer-internal work, Palmyra post-training, invented metrics, or safety research.
Honest fit boundary
Enterprise writing-product and Writer platform research is a stretch. I have not post-trained Palmyra or other Writer models, run RLHF, RLAIF, DPO, or GRPO at production LLM scale, designed Writer-internal evaluation benchmarks, or claimed Writer-internal work, and I do not invent metrics or safety research. The credible contribution is evaluation under uncertainty, production agent evaluation harnesses, scientific/HPC rigor, and data pipelines.
Role and location
AI research scientist · San Francisco, CA · Hybrid. Ashby lists workplaceType Hybrid and isRemote true, with secondary locations Seattle, WA and New York City, NY. The posting states this role is hybrid, based out of the San Francisco or New York City hub. Onsite-day count is not published. Willing to relocate to San Francisco or New York City with a relocation package. Fully remote work is not asserted.
Posting compensation: “$234.3K – $349K • Offers Equity”. · Official role posting