Skip to main content
Somesh PanchalSP
Looking for a job

Somesh Panchal

@someshpanchal

I optimize production LLM systems, cutting training time, inference latency, and GPU memory use.

United States
Message

At Qorvo, I optimize production LLM systems across multi-GPU Linux environments, model serving APIs, and performance monitoring. I reduced per-epoch training time by 35%, production inference latency by 40%, and GPU memory utilization by 25%.

I build reproducible ML services with PyTorch, FastAPI, Docker, MLflow, GitHub Actions, SQL, Parquet, and FAISS. My work spans distributed training, quantization, KV-cache reuse, request batching, streaming, caching, benchmarking, and production validation.

I also built an LLM Inference Optimization Benchmark Suite that achieved up to 3x throughput improvement and 60% GPU memory reduction, plus a multi-agent SQL analyst that turns business questions into validated SQL and reports. Earlier at DXC Technologies, I developed Python and SQL data pipelines, feature-engineering workflows, supervised-learning models, and automated data-quality monitoring.

Experience

Work history, roles, and key accomplishments

Education

Degrees, certifications, and relevant coursework

Southern Arkansas University logoSU

Southern Arkansas University

Master of Science, Computer and Information Science

2023 - 2025

Pursued a Master of Science in Computer and Information Science, focusing on advanced computing topics.

Gujarat Technological University logoGU

Gujarat Technological University

Bachelor of Engineering, Computer Engineering

2018 - 2022

Earned a Bachelor of Engineering in Computer Engineering, building a foundation in software and hardware systems.

Availability

Looking for a job

Location

United States

Authorized to work in

Salary expectations

80k+ USD

Skills

PythonCC++SQLBASHAlgorithmsData StructuresSoftware DesignSoftware ImplementationDebuggingVersion ControlGitCode ReviewTechnical DocumentationPyTorchHugging Face TransformersSentence Transformersscikit learnXGBoostLightGBMTensorFlowSupervised LearningClassificationFeature EngineeringNeural NetworksModel EvaluationStatistical ModelingLLM Fine tuningLoRAsQLoRAPEFTPrompt EngineeringRAGEmbeddingsSemantic SearchNatural Language ProcessingAI AgentsTool UseLangGraphLang ChainAI Driven AutomationPTQQATGPTQAWQBitsandbytesKV CachingFlash AttentionTorch.CompileTorchScriptONNX RuntimeTensorRTPruningKnowledge DistillationQuantizationBenchmarkingTesting and EvaluationSystem Performance TuningQuality Performance AnalysisDistributed Training (PyTorch DDP)FsdpDeepSpeedMixed PrecisionGradient AccumulationGradient CheckpointingPipeline ParallelismTensor ParallelismMulti GPU SystemsVLLMTriton Inference ServerTorchServefastAPIREST APIsRequest BatchingStreamingCachingAutoScalingApplication DeploymentSystem DeploymentProduction MonitoringProduction TroubleshootingFaissPgvectorPostgreSQLMySQLSQL ServerParquetJSONPandasNumPyData PipelinesBatch ProcessingData StorageData Quality ManagementSchema ValidationQuery ValidationDockerDocker ComposeContainerizationGitHub ActionsCI CDMLFlowModel RegistryPrometheusGrafanaLoggingObservabilityCloud MonitoringTesting AutomationModel VersioningAWSGoogle Cloud PlatformAzureCloud ComputingCloud InfrastructureNVIDIACUDALinuxAgileScrumSDLCSystem DesignProcess DesignEngineering ValidationSoftware TestingQuality AssuranceA B TestingFailure AnalysisRisk MitigationContinuous ImprovementProblem SolvingCross Functional CommunicationStakeholder CommunicationProject DeliveryStreamlitPyTorch ProfilerNVIDIA NsightNVIDIA SMI

Interested in hiring Somesh?

You can contact Somesh and 90k+ other talented remote workers on Himalayas.

Message Somesh

People also viewed

View all talent

Get matched with your dream remote job

Sign up now and join over 250,000+ remote workers who receive personalized job alerts, curated job matches, and more for free!

Sign up
Himalayas profile for an example user named Frankie Sullivan