Senior Software Engineer
Job Requirements
Advanced Natural Language Processing & Information Extraction
Experience with SOTA frameworks such as Thinc,Flair,zero-shot architectures like GLiNER. Knowledge of embedding generation, vector spaces,modern embedding models (e.g., google/embeddinggemma-300m).
Deep Learning & Model Optimization, High-Performance ML Operations & Backend Engineering
Understanding of the GIL,multi-threading,and memory/thread safety when serving heavy ML models in production web servers.
MLOps,Cloud & Data Lifecycle
Using MLflow for experiment tracking,model registry,reproducibility., Python 3.7+
FastAPI,Uvicorn,Starlette,Pydantic v2+pydantic-settings
pytest,pytest-asyncio(unit+functional tests)
asyncio,aiohttp, Project packaging(setuptools+pyproject.toml), REST API development
PyTorch 2.x(CPU/GPU/MPS), CUDA 12.*+, PyTorch Lightning, Hugging Face Transformers
Sentence_transformers
Classical ML,scikit-learn,NumPy,Pandas,SciPy,matplotlib
JupyterLab — exploration and training notebooks
Basic knowledge of Linear Algebra
Custom NLP pipeline design (spaCy,incl.transformer-encoder pipelines)
Custom spaCy components: NER,SpanCat,Dependency parser,Sentencizer,SpanFinder,custom tokenizers/matchers
Thinc,Flair NLP,GLiNER, Retrieval and embedding models(e.g.,google/embeddinggemma-300m)
Data processing, spaCy Projects, Prodigy, Text processing
Fuzzy search, Regex, XML processing, LLM, LLM-assisted data annotation with quality guards
LLM inference servers: vLLM,Hugging Face TGI,llama.cpp
Async Python (AsyncOpenAIClient) for high-throughput dataset processing
Cloud & integration, Docker, GitHub Actions
GCP: Artifact Registry,GCS,Compute Engine,GKE,Logging,Vertex AI,Gemini API, Jenkins,K8s,Dynatrace,ELK
Apache Kafka, MLflow, Git, NER and span classification
Sequence labeling and document-level classification,Transfer learning/fine-tuning transformers, Train/eval/deploy lifecycle
Data augmentation,information extraction,relation extraction,entity linking,information retrieval
Fuzzy matching,clustering (DBSCAN)
Precision/recall tradeoffs in information extraction
Thread safety serving models in multi-threaded web servers
XML processing; basic knowledge of doc(x),LaTeX,PDF formats
Academic publishing domain — manuscript structure,metadata standards
Other tech as a plus ,Java, JavaScript"
Work Experience
5-7Years